← Back to Table of Contents

Chapter 38 β€” The Frontier β€” Open Problems & What’s Next

β€œThe most exciting phrase in science is not β€˜Eureka!’ but β€˜That’s funny…’” β€” Isaac Asimov

This final chapter surveys the frontiers of language modelling β€” where rapid progress is being made, and what remains stubbornly unsolved.

Long Context

The push toward million-token context windows:

Model Max Context Method
GPT-4 Turbo 128K Proprietary
Claude 3.5 200K Proprietary
Gemini 1.5 Pro 2M Ring Attention + RoPE scaling
LLaMA-3.1 128K Progressive RoPE ABF training
Yarn / LongRoPE Extensible NTK-aware interpolation
Mamba/SSMs Theoretically ∞ Fixed-size recurrent state

Open challenges:

  • Retrieval across long context: models struggle to use information buried in the middle (β€œlost in the middle” problem)
  • Cost: attention is O(TΒ²) β€” 1M tokens means 10ΒΉΒ² operations per layer
  • Evaluation: few benchmarks test genuine long-context reasoning (RULER, BABILong are early attempts)

Efficiency Frontiers

Making models smaller, faster, and cheaper without sacrificing quality:

Efficiency Research Directions
Architecture Efficiency
MoE (sparse activation), SSMs (linear complexity), hybrid models, speculative decoding, early exit.
Compression
Quantization to 2-4 bits, pruning (structured + unstructured), distillation, neural architecture search.
Training Efficiency
Better data curation (quality > quantity), curriculum learning, parameter-efficient fine-tuning, continual pre-training.
Inference Optimization
KV-cache compression, prefix caching, request batching, hardware-aware kernel design.

The trend: smaller models trained on more, better data. LLaMA-3-8B matches LLaMA-2-70B. Phi-3-mini (3.8B) matches LLaMA-2-13B.

Multimodal

The convergence toward unified models that see, hear, read, and generate across modalities:

Multimodal Evolution
Text only
GPT-3, LLaMA β€” language in, language out
Vision + Language
GPT-4V, LLaVA, Gemini β€” understand images + text
Audio + Vision + Language
GPT-4o, Gemini 2 β€” native speech, image understanding
Generation
DALL-E 3, Sora, Veo β€” generate images, video from text
Unified
Any modality in β†’ any modality out (emerging)

Open frontiers: video understanding at scale, spatial reasoning, audio-visual grounding, embodied AI.

World Models

Can language models learn a model of the physical world β€” not just statistical patterns in text?

  • Video prediction: Sora, Veo generate physically plausible video
  • Simulation: models that can predict what happens next in a 3D environment
  • Embodied AI: robots using LLMs/VLMs for planning and control (RT-2, Figure)
  • Open question: is next-token prediction on enough data sufficient to learn world models, or is a fundamentally different approach needed?

Safety & Alignment

Safety Research Areas
Interpretability
Understanding what's happening inside models. Sparse autoencoders, probing, mechanistic interpretability, circuit discovery.
Robustness
Jailbreak resistance, adversarial robustness, consistent behavior under distribution shift.
Governance
EU AI Act, export controls, model evaluation for dangerous capabilities, responsible release practices.

Key debates:

  • Open vs closed weights: open models enable research but also misuse
  • Scaling risk: do larger models have qualitatively new risks?
  • Alignment tax: how much capability do we sacrifice for safety?

What Remains Unsolved

Problem Status
Reliable reasoning Improving rapidly (o1, R1), but still fails on novel problems
Catastrophic forgetting Models lose capabilities when fine-tuned for new tasks
Continual learning Can’t efficiently learn from a stream of new data without retraining
Causal reasoning Models learn correlations, struggle with true causal understanding
Planning Multi-step planning with long horizons remains fragile
Grounding Language models don’t truly β€œunderstand” β€” they pattern match (or do they?)
Efficiency at the frontier Training a frontier model still costs $100M+ and takes months
Evaluation No benchmark fully captures what we mean by β€œintelligent”

The Open-Source Flywheel

The Virtuous Cycle
Meta/Mistral/DeepSeek release open weights
β†’
Community fine-tunes, evaluates, improves
β†’
Research papers discover new techniques (LoRA, DPO, etc.)
β†’
Techniques adopted by labs β†’ better open models β†’ repeat

Looking Forward

The pace of progress in language modelling is extraordinary. A few predictions that are likely to age poorly:

  1. Models will get smaller and better β€” the 8B model of 2026 will match the 70B model of 2024
  2. Hybrid architectures will win β€” attention + SSMs + MoE, not a single architecture
  3. Reasoning will be the key differentiator β€” raw knowledge can be retrieved, reasoning can’t
  4. Inference will matter more than training β€” test-time compute scaling is just beginning
  5. Agents will become the interface β€” models that can act, not just answer

You’ve reached the end of the guide. πŸŽ‰

If you’ve read this far, you now have a comprehensive understanding of language modelling β€” from the mathematics of attention to the engineering of distributed training, from tokenization to the frontier of AI research.

The field moves fast. The best way to keep up is to read papers, run experiments, and build things.


Appendices:

← Previous: Chapter 37 β€” Agents & Tool Use Β· Back to Table of Contents


Last updated: April 2026