← Back to the journal

signals · September 2026

September 2026 in AI: the month AI started looking like infrastructure

September’s defining signal was not one model release. It was the expansion of AI into persistent agents, engineering systems, sovereign compute, energy, security, and the physical world.

September 2026 in AI: the month AI started looking like infrastructure

September delivered another dense month of AI announcements: new models, broader agent platforms, more autonomous coding workflows, and enormous infrastructure commitments. But the durable story was not any single release. It was the way these developments began fitting together.

That shift changes the questions leaders need to ask. Which model is best becomes which model is sufficient for this task? A clever assistant becomes a persistent actor inside a workflow. A model deployment becomes an operating system of context, tools, identity, evaluation, compute, energy, and human accountability.

The model race is becoming a model portfolio

September brought another concentrated wave of releases and updates across Anthropic, OpenAI, Google, DeepSeek, Meta, Qwen, and others. AI Release Tracker counted 12 model releases during the month. The exact count matters less than the pattern: enterprises can no longer treat model selection as a one-time platform decision.

Comparison between an older application using one best available model and a task-aware model router selecting fast, balanced, or frontier models
The enterprise question is shifting from “Which model is best?” to “Which model is sufficient for this task, under these cost, latency, and risk constraints?”

A practical model portfolio routes extraction, classification, routine agent work, coding, and hard reasoning differently. The objective is not maximum capability on every call. It is the lowest-cost route that clears the quality, latency, privacy, and safety bar for the task.

Routing signalOperational question
Task complexityDoes the task require frontier reasoning or a smaller fast model?
Consequence and riskWhat is the cost of an incorrect answer or action?
LatencyIs the work interactive, asynchronous, or batch?
Data boundaryWhich providers and deployment regions are permitted?
Evidence from evaluationWhich model actually succeeds on this workload?

Agents are becoming persistent

Microsoft’s new Copilot combined Home, Code, and Autopilot. Microsoft described Autopilot as a persistent, proactive agent that keeps working when the user is away. That is a more consequential change than a larger chat window: an assistant waits for a request, while a persistent agent can monitor, interpret, plan, act, remember, and observe again.

Persistent enterprise agent loop connecting observation, understanding, planning, action, memory, and repeated monitoring to email, calendars, documents, Teams, applications, databases, and APIs
Persistence turns an assistant into an operating participant. That makes ownership, permissions, stop conditions, and auditability part of the product.

Persistence increases value only when it is bounded. Teams need to define the event that wakes the agent, the systems it may observe, the actions it may propose or execute, the state it may retain, the conditions that require approval, and the evidence it must leave behind.

AI coding is becoming multi-agent engineering

September’s software-development discussion moved beyond autocomplete. Agentic development now includes subagents, delegation, cross-session communication, persistent environments, model routing, isolated execution, and harnesses that coordinate implementation and verification. O’Reilly’s September review highlighted organizations building their own agents and harnesses closely integrated with their working environments.

Timeline from manual development and AI-assisted coding to coding agents and multi-agent engineering coordinated by a specification and orchestrator
The scarce skill moves from typing every implementation detail toward specification, architecture, decomposition, validation, integration, and judgment.

The developer does not disappear. The role moves outward: define the specification, shape the architecture, divide the work, establish acceptance criteria, inspect evidence, resolve ambiguity, and integrate the result. Multi-agent speed without a shared specification can simply produce parallel inconsistency.

Related: Spec-driven development as the control plane for AI coding agents ↗

Harness engineering is becoming a real layer of the stack

As model capability improves, more engineering attention is moving to the system surrounding the model: context, memory, tools, state, planning, routing, execution, verification, recovery, evaluation, and observability. O’Reilly’s definition is useful: the model provides capability; the harness gives that capability a controlled way to complete work.

Enterprise AI system containing context, memory, tools, business logic, evaluation, state, observability, routing, and a model interface spanning multiple model providers
A model is a component. The surrounding system determines how intelligence is grounded, routed, observed, recovered, and converted into business outcomes.

This separation is strategically important. If models become easier to substitute, durable advantage will come from the organization’s context, evaluations, workflows, controls, and operating knowledge—not from a thin dependency on one model endpoint.

Deep dive: Harness engineering for AI agents ↗

AI agents are reaching physical systems

General Robotics announced updates to GRID, its platform for robot training and deployment. The company says the platform automates work across calibration, simulation, perception debugging, control-loop correction, and skill transfer. Reported setup improvements are company claims and still require independent validation, but the architecture is the signal worth watching.

Comparison showing the same plan, act, observe, evaluate, and adapt loop applied to coding agents and physical AI systems
The same feedback loop that made coding agents useful is moving into robotics, where actions meet sensors, actuators, simulation, control, and real-world consequences.

Physical AI raises the stakes. The environment is no longer only a repository or browser. Failures can involve equipment, people, production lines, and safety. Evaluation must therefore include scenario coverage, simulation-to-reality gaps, safe states, human intervention, and the physical limits of the system.

AI infrastructure became a national strategy issue

Canada welcomed Bell Canada’s planned Saskatchewan expansion, describing up to 900 megawatts of new capacity, a path toward a 1.2-gigawatt provincial AI infrastructure hub, and capital investment of up to C$52.5 billion. The announcement framed the project explicitly around sovereign Canadian compute.

Comparison between an earlier AI strategy focused on models, data, and talent and a broader strategy that also includes compute, energy, data centres, networking, and sovereignty
AI strategy now reaches beyond models, data, and talent into the physical and political foundations that make AI possible at scale.

Compute location affects resilience, latency, regulation, procurement, security, and industrial policy. For enterprises and governments, infrastructure architecture is becoming part of AI strategy—not merely a backend concern delegated after the use case is chosen.

The AI race is becoming an energy race

The scale of AI infrastructure exposes the power system underneath it. On September 30, a U.S. Senate proposal addressing whether large electricity users should bear more of the incremental grid-upgrade costs failed to advance. Whatever the legislative path, the policy question is now unavoidable: who pays for the generation, transmission, cooling, and grid expansion required by new data-centre demand?

AI cost model expanding from cost per token to cost per inference and cost per agent task, supported by GPUs, data centres, electricity, and the grid
Cost per successful task is the useful application metric, but every task still rests on physical compute, facilities, electricity, and grid capacity.

Efficiency therefore has several layers: fewer unnecessary model calls, better routing and caching, more efficient accelerators, higher infrastructure utilization, better cooling, and cleaner, more reliable power. The true cost of AI is not captured by a token price alone.

AI chips are becoming geopolitical infrastructure

Huawei introduced the Atlas 960 SuperPoD computing cluster as China continues developing domestic alternatives under export restrictions. The announcement reinforces a two-layer competition: model providers compete on capability and applications, while infrastructure ecosystems compete on chips, memory, networking, power, cooling, manufacturing, and supply chains.

Two-layer AI race with model providers above a compute infrastructure layer of GPUs, accelerators, memory, interconnects, networking, power, cooling, and manufacturing
Models depend on compute. The AI race is increasingly also a race over industrial capacity, supply chains, energy, and resilience.

Agent security moved from theoretical to operational

A traditional model security review often focuses on input, model behaviour, and output. An agent introduces a longer chain: untrusted input, retrieved context, memory, planning, authorization, tools, external systems, and outcomes. Each transition creates a new place where data, authority, or intent can be manipulated.

Comparison of a traditional input-model-output AI security model with an agent security model spanning reasoning, retrieved context, memory, proposed actions, authorization, tools, external systems, and outcomes
Agent security is system security. Guardrails have to follow data and authority through every step, not stop at the model boundary.
  • Identity: know which human, service, and agent initiated the action.
  • Delegated authority: preserve the scope, expiry, and provenance of permission.
  • Tool controls: validate parameters and enforce policy at the execution boundary.
  • Sandboxing: contain code, files, credentials, networking, and dependencies.
  • Runtime monitoring: record decisions, actions, outcomes, anomalies, and recovery.
  • Audit and evaluation: prove whether the system behaved within its approved envelope.

Related: Why enterprises need an agent registry ↗

The agent sandbox is becoming infrastructure

O’Reilly’s September review highlighted Docker Sandboxes as one example of disposable environments designed for AI agents. The design pattern is more important than the product: start from a clean environment, grant only the files, tools, dependencies, credentials, and network access required for the task, keep the approved artifacts and evidence, then discard the environment.

Agent executing in a disposable sandbox containing files, code, tools, browser access, and dependencies, then preserving artifacts while discarding the environment
Ephemeral execution reduces state leakage and limits the blast radius while making agent runs more repeatable.

This is another sign that agent architecture is converging with distributed systems, cloud infrastructure, and security engineering. Capable agents do not only need prompts. They need governed runtime environments.

The model may be becoming less important than the system

Put these developments together and the centre of gravity moves. The model remains critical, but enterprise value increasingly depends on the system that supplies context, selects models, constrains tools, executes workflows, observes outcomes, and improves over time.

AI system architecture combining context, memory, data, models, routing, evaluation, tools, APIs, and sandboxes into workflows and outcomes
The production unit is no longer the model endpoint. It is the whole system that turns a business objective into a governed outcome.
AI stack expanding from models to agents, harnesses, sandboxes, tools, infrastructure, and energy
September’s developments form one connected stack: intelligence, action, orchestration, isolation, integration, compute, and power.

The bigger story from September

If September 2026 can be summarized in one progression, it is the movement from content generation to operational systems. The questions changed from “Can AI create something useful?” to “How do intelligence, software, infrastructure, energy, and people operate as one reliable system?”

Evolution from generative AI in 2023 through enterprise GenAI, agentic AI, integrated AI systems, and a future AI infrastructure ecosystem
The surface area of AI is expanding: from models to agents, agents to workflows, workflows to infrastructure, and digital intelligence into the physical world.

What I am watching next

Three signals for the months ahead

01Agent economics

Cost per successful task—including model calls, tools, retries, evaluation, infrastructure, and human review—will become more useful than cost per token.

02Physical AI

Harnesses, environments, feedback loops, and objective evaluation are moving from software agents into robotics and industrial systems.

03Infrastructure efficiency

Compute availability, inference efficiency, energy, networking, and sovereign capacity are becoming strategic design constraints.

Sources and further reading

PRIMARY SOURCES

Reporting and primary material used for this newsletter

Models, agents, and engineering

Models, agents, and engineering
O’Reilly — Own the outer loop

A practical view of agent engineering as model capability plus files, tools, memory, skills, sandboxes, permissions, observability, and recovery.

Physical systems and infrastructure