RNNs and LSTMs process sequences one token at a time, carrying a hidden state forward as memory.
ConceptWhat it is
A Recurrent Neural Network (RNN) reads a sequence one element at a time, updating a hidden state that is meant to summarize everything seen so far, so it exists to model order and time-dependence in data like text, speech, or sensor streams. The LSTM (Long Short-Term Memory) variant adds gates that control what to remember, forget, and output, fixing the plain RNN's tendency to lose long-range signal.
These architectures were the dominant approach to language and sequence modeling before the transformer showed that attention could do the same job faster and in parallel.
How it worksThe mechanics
At each timestep the network takes the current input and the previous hidden state, applies learned gates in the LSTM case to decide what old information to keep or discard, and produces a new hidden state and an output; because step t+1 depends on the output of step t, the computation is inherently sequential and cannot be parallelized across time the way attention can.
At a glanceSee it
Inside an LSTM cell — three gates plus a separate long-term cell state that lets memory flow across many steps almost unchanged.
Choosing among RNN, LSTM or GRU, and attention by how far back the dependency reaches — with the vanishing-gradient failure that dooms the plain RNN.
When to use itWhere it fits
- Small, streaming, or resource-constrained sequence tasks like keyword spotting on embedded devices.
- Time-series forecasting where sequence length is short to moderate.
- Legacy systems already built around recurrent architectures that are not worth migrating.
- Teaching the core concept of sequential memory before introducing attention.
When NOT to use itLimits & anti-patterns
- Long documents or long-range dependencies, where vanishing gradients and sequential bottlenecks hurt quality and training speed compared to transformers.
- Any task needing large-scale parallel training on modern GPUs, since sequential computation cannot be parallelized across time steps.
- State-of-the-art language generation, where transformer-based LLMs now outperform RNNs on essentially every benchmark.
Trade-offsAdvantages & costs
Advantages
- Naturally models order and variable-length sequences with a small memory footprint.
- LSTMs handle moderately long dependencies far better than plain RNNs.
- Cheap to run at inference for short sequences on constrained hardware.
- Conceptually simple and well understood after decades of use.
Trade-offs & costs
- Sequential computation prevents parallel training, making it slow at scale.
- Struggles with very long-range dependencies despite LSTM gating.
- Largely superseded by transformers for language and most sequence benchmarks.
- Harder to stack very deep compared to transformer blocks with residual connections.
ExampleIn the real world
Google's early production machine translation system, pre-Transformer, used stacked LSTM encoder-decoder networks to translate search queries and documents.ToolsHow to implement it
- PyTorch nn.LSTMstandard building block for recurrent models.
- Keras SimpleRNN/LSTM layersquick prototyping for time-series and text.
- ONNX Runtimedeploying trained RNNs to edge and mobile.
- Darts / statsmodelstime-series libraries that still lean on recurrent baselines.
Cost & effortWhat it takes
Cheap to train on modest hardware for short sequences; inference latency is low but does not parallelize across time, so throughput lags transformers at scale.