The Data Shows a Real Fleet-Management Cohort
The October 2026 AI Leaders AI Leverage Report asked respondents how many AI agents their organizations were running in production. Among the 75 people who answered that question, 60 reported at least one production agent. Twenty-two selected the highest range, more than 20.
| Production agents reported | Respondents | Share of base 75 |
|---|---|---|
| None | 9 | 12% |
| 1-5 | 27 | 36% |
| 6-20 | 11 | 15% |
| More than 20 | 22 | 29% |
| Not sure | 6 | 8% |
Source: AI Leaders AI Leverage Report, October 2026. Base 75; self-reported responses.
The distribution shows that operating a fleet is already a practical issue for some respondents. It does not establish a market average or prove that more agents produce better results. The survey captured self-reported counts, not independently verified uptime, autonomy, reliability, financial value or control quality. The separate answer on production-agent counts provides the concise result and its limits.
Why a Fleet Is More Than a Longer Inventory
With two or three agents, a technical lead may remember which prompt changed, which credential is shared and which business owner approved a release. That memory stops being a control as the fleet grows.
A fleet of more than 20 agents introduces dependencies that are easy to miss:
- Several agents may call the same model endpoint or retrieval service.
- One agent may trigger another agent, creating a multi-step failure path.
- A prompt, model or tool-policy update may affect several workflows at once.
- A shared data connector or credential can create a fleet-level blast radius.
- An agent may be technically active after its owner, purpose or downstream process has changed.
- A retired vendor endpoint can leave an apparently healthy workflow unable to complete.
This is why an inventory alone is insufficient. Inventory tells leaders that an agent exists. A fleet operating model tells them which version is live, what service it is expected to provide, what can affect it, how to contain it and how to remove it.
Detailed identity and permission controls belong in the security operating model. Open Future Forum covers those controls in What CISOs Are Doing About AI Agent Security and Governance and Shadow AI vs. AI Agent Risk. The framework below concentrates on operating a changing fleet after those basic controls exist.
1. Give Every Agent a Stable Identity and Lifecycle State
Create one agent identifier that does not change when the underlying model, vendor, prompt or deployment environment changes. A stable identifier makes release history, incidents, cost and retirement evidence traceable over time.
The live state should be explicit. A practical state model is:
| Lifecycle state | Meaning | Minimum exit condition |
|---|---|---|
| Proposed | A bounded use case is being considered | Named workflow, sponsor and initial risk tier |
| Development | The agent is being built and tested | Evaluation results and release candidate |
| Limited release | Use is restricted by users, volume or actions | Acceptance criteria met at limited scope |
| Full production | Approved users and workloads can depend on it | Service owner accepts ongoing obligations |
| Restricted | The agent remains live with reduced actions or audience | Remediation completed or replacement approved |
| Paused | Execution is stopped but evidence and dependencies remain | Restart approval or retirement decision |
| Retired | Execution, access and scheduled calls are removed | Closure evidence and retention obligations completed |
These are recommended operating states, not states measured in the survey. The important control is that a state has entry and exit criteria. “Pilot” should not become a permanent description for software that performs daily production work.
The record should also show the business workflow, business owner, technical owner, service tier, current release, downstream consumers and next review trigger. It should link to cost and control evidence without duplicating those systems.
2. Treat Prompts, Models, Tools and Data as Versioned Components
An agent can change materially without a new application release. A provider can update a model. A team can edit a system prompt. A new retrieval collection can alter the information available to the agent. A tool-policy change can let it take an action it could not take yesterday.
For every production release, record at least:
- The stable agent identifier and release number.
- Model provider, model name and version or dated endpoint where available.
- Prompt or instruction-set version.
- Enabled tools, action limits and approval points.
- Retrieval source or knowledge-release version.
- Evaluation pack and acceptance result.
- Change ticket, approving owner and release time.
- Rollback target and person authorized to invoke it.
Define which changes are material. A change is normally material when it can alter the agent’s output, available information, permitted actions, cost profile or failure mode. Material changes should pass a stated release path, even when the vendor labels the update as routine.
At fleet scale, change control should also identify correlated exposure. If 14 agents depend on the same model endpoint, a model migration is not 14 unrelated changes. It is one fleet event with 14 workflow consequences. Release in stages, compare evaluation results, set a rollback point and watch aggregate service signals before completing the migration.
3. Give Each Agent a Service Tier
“In production” does not state how reliable an agent must be or how quickly a failure must be addressed. Assign a service tier based on workflow consequence and dependency, not the excitement surrounding the use case.
| Illustrative tier | Typical use | Evidence to define |
|---|---|---|
| Tier 1: workflow-critical | An interruption stops or materially delays an important process | Availability window, completion target, maximum interruption, fallback and incident response owner |
| Tier 2: material support | An interruption reduces capacity or quality but a manual path remains | Completion and escalation targets, manual capacity, recovery objective and review owner |
| Tier 3: assistive | Users can continue without the agent | Basic health signal, support route and retirement trigger |
The table is a recommended classification, not a surveyed practice. Each organization should set its own targets. Useful evidence can include completion rate, unacceptable-output rate, human escalation rate, latency, cost per completed unit and time spent in fallback. A single fleet-wide uptime target can hide very different business consequences.
The service record should define the unit being measured. An agent can return a response while the workflow still fails. For a contract-review workflow, the relevant unit might be a review completed with required fields and citations, not a successful model call.
4. Build Containment at Agent, Workflow and Fleet Level
Stopping one agent is not always enough. The same credential, connector, model, retrieval source or orchestration service may be used elsewhere. A fleet operating model needs several containment levels:
- Agent containment: stop one agent release or disable one tool.
- Workflow containment: stop the complete process, including triggers and downstream actions.
- Dependency containment: revoke a connector, credential, data source or model endpoint used by several agents.
- Fleet containment: pause a class of agents affected by the same incident or unapproved change.
Record dependency groups before an incident. Useful groupings include provider, model version, credential, data connector, tool, orchestration platform and high-consequence action. This turns an incident question from “Which agents might be affected?” into a reviewable list.
Containment should preserve evidence. The stop action should prevent further execution while retaining the release record, relevant logs, decision history and owner trail according to the organization’s retention rules. A kill switch without an evidence-preservation plan can stop the immediate problem and weaken the later investigation.
Fleet-level reporting to the board should remain compact. The AI Dashboard Every Board Should Request From Management explains the board view. The operating team needs the deeper dependency and release evidence described here.
5. Separate the Fleet Record From Workflow Economics
The fleet record answers, “What is live, what version is it, and can the organization trace and control it?” The cost record answers, “What does the workflow consume?” The value case answers, “Is it worth continuing?” Keep these linked, but do not collapse them into one status label.
For each agent, point to a workflow identifier and cost object. Where several agents complete one workflow, the workflow should carry the business unit cost while the fleet record retains agent-level usage and service evidence. This avoids treating a low model invoice as proof of an efficient end-to-end process.
The October 2026 CFO AI Leverage Report recommends tracking model, infrastructure, data, integration, monitoring, security, legal and human-review costs together. Detailed allocation mechanics belong in a workflow-level cost-accounting guide rather than inside the fleet register. These are management frameworks, not practices measured by the survey.
6. Make Retirement a Controlled Production Change
Retirement is not deleting a row from the register. It is a production change that can break scheduled jobs, downstream reports, customer communications or another agent’s tool chain.
A retirement decision can be triggered by:
- No current business or technical owner.
- Persistent failure to meet the service expectation.
- Duplicate capability after a platform consolidation.
- Expired data rights, contract terms or credentials.
- A workflow redesign that removes the original need.
- Cost or control evidence that no longer supports operation.
- Replacement by a new release, agent or non-AI process.
Before retirement, map incoming triggers and downstream consumers. Notify affected owners. Remove schedules and event triggers. Revoke credentials and tool access. Confirm how records, prompts, evaluations and logs should be retained. Observe the workflow after removal, and document that no unknown dependency attempted to call the retired service.
If the agent is replaced, keep the old and new release identities distinct. A replacement should not erase the incident, cost or service history of the system it superseded.
A Practical Fleet Review Cadence
The survey did not measure review frequency. A useful operating rhythm combines event-driven reviews with a recurring fleet review.
Review an individual agent after a material model, prompt, tool, data, credential or workflow change. Review dependency groups after a provider or platform change. Then review fleet exceptions on a recurring schedule.
The recurring meeting should focus on decisions:
- Which agents changed state or version since the last review?
- Which releases missed their service expectations?
- Which dependencies now create concentrated exposure?
- Which agents lack a current owner, rollback target or tested fallback?
- Which paused or duplicate agents should be retired?
- Which exceptions need remediation, restriction or fleet-level containment?
The output is an exception list with owners and dates. A complete register with no decision trail is administration, not fleet control.
Key Citable Facts
- 22 of 75 respondents (29%) reported more than 20 production AI agents in the October AI-leader instrument.
- 60 of 75 respondents (80%) reported at least one production AI agent.
- 29 of 75 respondents (39%, any mention) mentioned data access and quality as a production bottleneck.
- Among 49 applicable responses to the access-method question, 24 (49%) reported per-agent credentials and 18 (37%) reported shared service accounts.
For reusable wording, denominators and source notes, use the October 2026 Executive AI Statistics rather than copying a percentage without its base.
Methodology and Caveat
These findings come from first-party questions embedded in the Enterprise AI at Microsoft application flow. The full export contained 919 rows, including invited-only records. The analysis excluded those records and used 314 eligible non-invited applications. Responses were deduplicated by email for each question using the latest available answer. Item bases range from 71 to 101 because questions were added or completed at different points in the flow.
The data is selective and self-reported. It should not be treated as representative of all companies, all executives or the wider market. The analysis is descriptive and does not establish causality. The lifecycle, service-tier, containment and retirement controls in this article are recommendations informed by the operating questions raised by the data. They were not measured as adopted practices.
Last updated: October 3, 2026
Frequently Asked Questions
Discuss Production AI With Peers
Open Future Forum convenes AI and operating leaders in small, off-the-record gatherings. If you are building the operating model for a production AI fleet, apply to join an Open Future Forum gathering.