This is the last piece in a twelve-week series, and I want to close it on something forward-looking rather than another pattern that's already shipped. So this post is about two things that are still nascent — world models and continual learning at the harness layer — and what their convergence tells us about where this whole thing is heading.
The thread I've been pulling on since Week 1 is that the bottleneck keeps migrating outward. Prompts to contexts. Contexts to harnesses. Harnesses to meshes. Meshes to governance. The forward direction from here is the one I want to spend this post on: from operational systems to learning systems. From agents that are good today to agents that are getting better every week, every quarter, every year — and not because the model weights change, but because the harness around them gets smarter.
World models, briefly
In April, MIT Technology Review ran a piece on the world-models race. The framing they used is right: language models are getting brittle on tasks that require an internal model of the physical world. A model trained on New York taxi data can give directions in New York — until you ask it for a detour, at which point it falls apart. It does not have a model of the city. It has a memorization of the city.
World models are the attempt to give AI systems internal representations they can simulate against. Three pieces of news converged in 2025-2026 to make this a real category:
DeepMind's Genie 3. A general-purpose world model that generates real-time interactive 3D environments from text prompts. 24 frames per second, 720p, with object permanence across scene transitions. Not a video — a navigable world.
World Labs' Marble. Fei-Fei Li's company shipped its first commercial product in November 2025 — Marble — which generates persistent, downloadable 3D environments from text, images, video, or panoramas. Uses 3D Gaussian Splatting. Available now in freemium and paid tiers.
Yann LeCun's new company. LeCun left Meta in 2026 to start a world-model-focused company. The thesis, consistent with what he's argued for years: large language models are insufficient for general intelligence; you need models that learn from observation of the physical world, not just text.
The interesting question is not "are world models real" — they are. The interesting question is what happens when they bleed out of robotics and gaming, where they are currently most useful, into mainstream product surfaces.
Where world models matter for non-robotics products
The applications most people miss are the boring ones:
Design. Architects and product designers using world models to generate, navigate, and test spatial designs before building. CAD with a navigable simulation attached.
Planning. Supply chains and logistics using world models to run scenarios. Not flat optimization — actual simulation, with the physical constraints baked in.
Forecasting. Financial and operational forecasting that simulates from a causal model rather than fitting curves to historical data. The model can answer counterfactuals.
Training environments. Customer service agents trained against simulated customer environments rather than scripted scenarios. The simulation has more variance, the agent generalizes better.
Each of these is a multi-year shift, and I am not predicting that any of them are dominant in 2026. What I am saying is that the foundation models are now good enough that these applications are starting to show up in serious product roadmaps. Three years out, the design tools, planning tools, and forecasting tools you use will have world-model components in them. Probably without making a big deal about it. The "AI" label will fade as the capability becomes routine.
Continual learning at three layers
The second half of this post is about how systems improve. Harrison Chase's April 5 piece on continual learning made the most useful taxonomy I've seen in 2026:
Model layer. Weights change. Expensive, slow, hard to control, hard to attribute improvements to specific changes.
Harness layer. The code wrapping the agent changes. Better tools, better verification, better guardrails. Medium cost, faster cycle.
Context layer. The instructions, skills, and tools the agent reads. The CLAUDE.md changes. The new skill gets added. Cheap, fast, immediately deployable.
Most of the visible AI progress in the public eye is at the model layer — new releases, new frontier scores. Most of the useful progress inside production systems is at the harness and context layers. And the gap is widening.
This is why the example Chase highlights — OpenClaw's SOUL.md, which updates over time as the agent learns — is the early signal worth paying attention to. The SOUL.md is a manifest file that the agent itself proposes updates to, based on what it learned in the previous run. The human reviews the proposed update. Approved updates land in the file. Next run, the agent is meaningfully different — without any weight change.
The shape of this loop is what I think the next era of agent systems looks like. Not magic. Not model upgrades. The agent's manifest grows from operation.
The Hashimoto loop, generalized
Compare this to Mitchell Hashimoto's principle from Week 7: every time the agent makes a mistake, you engineer a fix so it never makes that mistake again. Today, in 2026, the "you" in that sentence is a human engineer. Tomorrow, in some bounded form, the "you" is the agent itself proposing the fix and a human approving it.
That is continual learning at the context layer. It is not the agent rewriting its own weights. It is the agent proposing changes to the artifacts it depends on, with a human in the loop, on a cadence the system supports.
The reason this is the closer for this series is that this is the operating model that takes everything we've built — spec-driven plans, context engineering, MCP, capabilities, orchestration, manifests, harnesses, evals, identity, mesh, governance — and turns it into a system that gets better every day, without a deployment hot fix, without a model release, without anyone in the room being heroic.
The end of "good today, mediocre next quarter"
Here is the failure mode I've watched at dozens of organizations through 2024-2025. They stand up an agent program. It works well for the first six months. Then the world drifts, the model gets updated, the team changes, and the system slowly degrades. By month 12, performance is meaningfully worse than at launch. By month 18, the system is being rebuilt.
Continual learning at the harness level is what breaks that cycle. The system isn't static. The CLAUDE.md is updated weekly. The skills library grows. The eval set adds the new failure modes you discovered. The harness encodes everything you've learned about what works in your business.
Twelve months in, the system is better than it was at launch, not worse. Twenty-four months in, the system is substantially better — but in a way that is mostly invisible from the outside, because the improvements are in the harness, the manifest, the eval set. The thing the customer sees just keeps working a little more reliably.
That is the operating posture I am pushing every client toward. Not "ship a magical AI thing." Ship a system that learns at organizational speed, in artifacts you control.
Closing the series
I'll close this twelve-week series on the same note I closed the London talk:
The technical layers commoditize. Prompts. Contexts. Harnesses. Meshes. Each becomes shared, standardized, available to anyone who can read the spec and pay for the infrastructure. The competitive moat keeps moving up the stack.
What does not commoditize is judgment. What to build. Who to build it for. How to govern it once it is running. Which experiments to run, which experiments to kill, what your organization is actually for. Those questions sit above all of the technical layers, and no amount of model capability replaces having a clear answer to them.
If there is one thing I want every reader to take from this series, it is this. The competitive question for your business in 2027 is not whether you adopted AI. By 2027 everyone will have. The competitive question is whether you built the discipline — spec-driven, harnessed, evaluated, identified, governed, continually learning — that turns adoption into compounding advantage.
That discipline is what my team at Applied Futures spends every day building with companies. If any of the twelve posts in this series resonated, that is the work we do — embedded engagements, bootcamps, workshops, hackathons. Practitioner-led. Hands-on. The opposite of slide decks.
Thanks for reading. The conversation continues — on LinkedIn, in person at meetups in Copenhagen, and on the next stage I'm on.

About the Author
Jacob Langvad Nilsson
Technology & Innovation Lead
Jacob Langvad Nilsson is a Digital Transformation Leader with 15+ years of experience orchestrating complex change initiatives. He helps organizations bridge strategy, technology, and people to drive meaningful digital change. With expertise in AI implementation, strategic foresight, and innovation methodologies, Jacob guides global organizations and government agencies through their transformation journeys. His approach combines futures research with practical execution, helping leaders navigate emerging technologies while building adaptive, human-centered organizations. Currently focused on AI adoption strategies and digital innovation, he transforms today's challenges into tomorrow's competitive advantages.
Ready to Transform Your Organization?
Let's discuss how these strategies can be applied to your specific challenges and goals.
Get in touchRelated Services
Related Insights
Agentic Mesh and the Governance That Has to Come With It
On August 2 the EU AI Act's main provisions became enforceable. The honest story isn't that compliance became a legal review — it's that compliance became an architectural constraint. The architecture that responds well is the agentic mesh. Here's what that actually looks like, and the order to build it in.
The Hybrid Wins: Small Language Models in Real Products
A 3.8B parameter model now matches GPT-4o on extraction tasks. Phi-4, Gemma 4, and Qwen 3.5 changed the calculus for what runs locally and what runs in the cloud. The shift isn't 'small models won' — it's hybrid as the default architecture. Here's the pattern that actually ships.