Chapter 38 β The Frontier β Open Problems & Whatβs Next
βThe most exciting phrase in science is not βEureka!β but βThatβs funnyβ¦ββ β Isaac Asimov
This final chapter surveys the frontiers of language modelling β where rapid progress is being made, and what remains stubbornly unsolved.
Long Context
The push toward million-token context windows:
| Model | Max Context | Method |
|---|---|---|
| GPT-4 Turbo | 128K | Proprietary |
| Claude 3.5 | 200K | Proprietary |
| Gemini 1.5 Pro | 2M | Ring Attention + RoPE scaling |
| LLaMA-3.1 | 128K | Progressive RoPE ABF training |
| Yarn / LongRoPE | Extensible | NTK-aware interpolation |
| Mamba/SSMs | Theoretically β | Fixed-size recurrent state |
Open challenges:
- Retrieval across long context: models struggle to use information buried in the middle (βlost in the middleβ problem)
- Cost: attention is O(TΒ²) β 1M tokens means 10ΒΉΒ² operations per layer
- Evaluation: few benchmarks test genuine long-context reasoning (RULER, BABILong are early attempts)
Efficiency Frontiers
Making models smaller, faster, and cheaper without sacrificing quality:
The trend: smaller models trained on more, better data. LLaMA-3-8B matches LLaMA-2-70B. Phi-3-mini (3.8B) matches LLaMA-2-13B.
Multimodal
The convergence toward unified models that see, hear, read, and generate across modalities:
Open frontiers: video understanding at scale, spatial reasoning, audio-visual grounding, embodied AI.
World Models
Can language models learn a model of the physical world β not just statistical patterns in text?
- Video prediction: Sora, Veo generate physically plausible video
- Simulation: models that can predict what happens next in a 3D environment
- Embodied AI: robots using LLMs/VLMs for planning and control (RT-2, Figure)
- Open question: is next-token prediction on enough data sufficient to learn world models, or is a fundamentally different approach needed?
Safety & Alignment
Key debates:
- Open vs closed weights: open models enable research but also misuse
- Scaling risk: do larger models have qualitatively new risks?
- Alignment tax: how much capability do we sacrifice for safety?
What Remains Unsolved
| Problem | Status |
|---|---|
| Reliable reasoning | Improving rapidly (o1, R1), but still fails on novel problems |
| Catastrophic forgetting | Models lose capabilities when fine-tuned for new tasks |
| Continual learning | Canβt efficiently learn from a stream of new data without retraining |
| Causal reasoning | Models learn correlations, struggle with true causal understanding |
| Planning | Multi-step planning with long horizons remains fragile |
| Grounding | Language models donβt truly βunderstandβ β they pattern match (or do they?) |
| Efficiency at the frontier | Training a frontier model still costs $100M+ and takes months |
| Evaluation | No benchmark fully captures what we mean by βintelligentβ |
The Open-Source Flywheel
Looking Forward
The pace of progress in language modelling is extraordinary. A few predictions that are likely to age poorly:
- Models will get smaller and better β the 8B model of 2026 will match the 70B model of 2024
- Hybrid architectures will win β attention + SSMs + MoE, not a single architecture
- Reasoning will be the key differentiator β raw knowledge can be retrieved, reasoning canβt
- Inference will matter more than training β test-time compute scaling is just beginning
- Agents will become the interface β models that can act, not just answer
Youβve reached the end of the guide. π
If youβve read this far, you now have a comprehensive understanding of language modelling β from the mathematics of attention to the engineering of distributed training, from tokenization to the frontier of AI research.
The field moves fast. The best way to keep up is to read papers, run experiments, and build things.
Appendices:
β Previous: Chapter 37 β Agents & Tool Use Β· Back to Table of Contents
Last updated: April 2026