Agents everywhere, adoption somewhere.
In the eighteen months since we wrote about rewiring the audit and consulting value chain with generative and agentic AI, the conversation has moved fast. Every audit technology vendor now describes its roadmap in agentic terms. Firms have announced agent platforms, agent marketplaces, and agentic audit visions. The ambition is right. The audit profession, squeezed between rising regulatory expectation and a persistent talent shortage, needs leverage of exactly this kind.
Yet inside engagement teams, the day-to-day reality has changed less than the announcements suggest. Auditors use AI to draft, summarise, and search. Few engagements run AI through the substance of the work — risk assessment, testing, evidence evaluation — in any systematic way. The gap between the agentic narrative and the audit file is wider than most of us would like, and it is worth exploring why.
Our view — offered as one practitioner’s perspective rather than the last word — is that the conversation has too often started at the wrong end. It begins with the technology — what can agents do? — and works backwards to the audit. We believe it should begin with the work.
‘Agentic washing’ is the rebadging of existing automation, scripts, and rule-based workflows as autonomous AI agents. A scheduled reconciliation macro becomes a ‘reconciliation agent’. A document comparison feature becomes an ‘agentic review capability’. The label changes; the work does not.
This is rarely deliberate. Vendors and firms are responding to real demand, and the boundary between sophisticated automation and genuine agency is honestly blurry — reasonable people draw the line in different places. We would simply suggest that audit, of all professions, has good reason to value precision here. The profession’s value rests on professional scepticism — the discipline of not accepting an assertion without testing it — and there is something fitting in extending that same scepticism to claims about our own tooling.
The labelling also matters practically: a true agent that plans, acts, and adapts needs supervision structures, guardrails, and documentation that a deterministic workflow does not. If everything is called an agent, governance can drift in either direction — over-engineered for simple automation, or too light for genuine autonomy that happens to sit in a bucket labelled ‘just another tool’.
The antidote to agentic washing is precision about work.
How do organisations determine which audit tasks are best suited for AI agents, assistants, or traditional automation?
Strip away the terminology and AI in the enterprise is about one thing: how work gets done, and by whom — or by what. An audit is a useful place to apply that lens because, despite its complexity, it decomposes cleanly. A statutory audit is not a monolith. It is a few hundred discrete tasks: ingest the trial balance, map the chart of accounts, assess risk by assertion, design procedures, send confirmations, test journal entries, evaluate estimates, scan subsequent events, draft the report, form the opinion.
Each of those tasks has a different shape. Some are deterministic and high-volume. Some require synthesis across unstructured documents. Some demand judgement that standards explicitly reserve for the auditor. Treating them as one undifferentiated mass — ‘the audit’ — and asking ‘how do we apply agents to it?’ produces vague answers and stalled pilots. Decomposing the work and asking, for each task, ‘what kind of intervention does this need?’ produces a deployable plan.
Crucially, the answer to that question is not always an AI agent. For any given task there are at least four honest answers: an AI agent, a copilot or assistant working alongside the auditor, an existing point solution that already does the job well, or a capability that should come from the core platforms — the audit software, the client’s ERP, the firm’s knowledge systems. Choosing ‘no agent needed’ is a legitimate and often correct design decision. It is also the decision that is easiest to lose sight of amid the current enthusiasm.
A spectrum, not a switch: the ‘Human + AI’ Service Autonomy Model.
To make these task-level decisions consistently, we use TCS’s ‘Human + AI’ Service Autonomy Model — a five-level spectrum describing how work and decisions are executed as AI capability deepens:
Two properties of the model matter for audit. First, decision-making shifts gradually from human to human-plus-AI as you move along the spectrum — there is no cliff edge where the auditor disappears. Second, the levels are assigned to tasks, not to the audit as a whole. A single engagement will, by design, run different tasks at different levels simultaneously. That is the point. The model turns ‘are we doing agentic audit?’ — a question that invites vague answers — into ‘which tasks sit at which level, and why?’ — a question that invites evidence.
How can firms build a scalable operating model for AI-enabled audits using task-level autonomy?
Applied to a typical engagement, the placement exercise might look something like the table below. It is illustrative rather than prescriptive — every firm’s methodology, tooling, and risk appetite will shift the placements — but the shape of the thinking is what matters:
Audit task |
Right intervention |
Why |
Trial balance ingestion and account mapping |
Core platform capability |
Structured, repeatable, deterministic. This belongs inside the audit platform, not bolted on as an agent. |
External confirmations |
Existing point solution |
Mature platforms already handle this with established controls and audit trails. Nothing to reinvent. |
Drafting engagement letters, prepared by client requests, and workpaper narratives |
Assistant (Level 2) |
The auditor stays in the driving seat; AI removes drafting effort, not judgement. |
Journal entry testing across the full population |
Supervised agent (Level 3) |
An agent can test 100% of entries against risk criteria; the auditor reviews exceptions and sets the criteria. |
Subsequent events and contradictory evidence scanning |
Supervised agent (Level 3) |
Continuous scanning of filings, news, and client data, with findings routed to the engagement team for evaluation. |
Analytical review of revenue against expectations |
Assistant, moving to supervised agent (Levels 2–3) |
An agent can build the expectation from operational and market data and flag variances; the auditor decides which variances matter and what they mean. |
Contract and lease review for accounting treatment |
Supervised agent (Level 3) |
An agent can extract key terms across a large contract population and propose a treatment; significant or finely balanced judgements go to the auditor. |
Continuous monitoring of recurring, low-risk control populations |
Autonomous AI workforce (Level 4), selectively |
Where criteria are stable and risk is low, agents can run end to end, with humans setting goals and sampling outcomes. |
Materiality, fraud risk assessment, going concern, the opinion |
Human-led, AI as tool (Levels 1–2) |
Professional judgement and scepticism under ISA 200 are non-delegable. AI informs; the auditor decides. |
Three things tend to stand out when teams run this exercise in earnest. First, a meaningful share of tasks land on ‘existing point solution’ or ‘core platform capability’ — the less celebrated answers. Second, the highest near-term value often concentrates at Level 3, where full-population testing and continuous evidence scanning expand audit coverage in ways sampling never could. Third, the tasks that define the profession — scepticism, judgement, the opinion — stay firmly human, and the model makes that explicit rather than leaving it to anxiety or accident.
What are the limits of AI autonomy in audit, risk assessment, and regulatory compliance?
It is worth slowing down on a single task, because the discipline only becomes real at that resolution. Take journal entry testing — a standard, mandated procedure aimed at management override of controls. Decomposed, it is not one task but several. Extracting the full journal population from the ledger and normalising it is a core platform job: structured, deterministic, and best handled by the audit software or a mature data tool, not an ‘agent’.
Scoring every entry against risk criteria — round-sum amounts, unusual postings to revenue, entries booked by unexpected users or at odd hours, descriptions that read oddly — is a supervised-agent job, where the agent works across 100% of the population rather than a sample. Setting those criteria in the first place, and deciding which flagged entries genuinely warrant investigation, is human work that the standards reserve for the auditor. And writing up the rationale for the entries selected and the conclusions reached sits comfortably with an assistant, with the auditor signing the words.
So a procedure many would label, in a sentence, as a candidate for ‘an agent’ turns out on inspection to involve four different interventions at three different levels — and one of them is a platform capability with no agent at all.
The same exercise repays the effort on: estimates, where an agent might assemble the comparable data and recompute a range while the auditor challenges management’s assumptions; on going concern, where an agent can monitor covenant headroom and cash runway continuously while the conclusion stays firmly with the engagement team; and on group audits, where much of the component-to-group reconciliation is structured work that need not consume senior time at all. None of this is a criticism of the ‘agent’ framing — it is simply what the framing looks like once the work is laid out honestly.
It is worth being direct about the right-hand end of the spectrum. Levels 4 and 5 will arrive in audit, but in pockets — continuous controls monitoring, recurring low-risk populations, internal firm operations — not at the opinion. ISA 200 places professional judgement and professional scepticism at the centre of the audit, and responsibility for the opinion with the engagement partner. That responsibility is non-delegable, to a junior or to an agent. Regulators have signalled the same expectation: AI can change how evidence is obtained and evaluated, but accountability does not move.
Far from limiting the opportunity, this constraint clarifies it. The economic prize is not removing the auditor; it is redirecting scarce, expensive professional judgement away from tasks machines now do better — and towards the risk assessment, estimates, and entity-level questions where judgement actually earns its fee. This is what we mean by auditors operating at the top of their license.
Task-level placement answers ‘what intervention does each task need?
Task-level placement does not, on its own, answer ‘how does this hold together at firm scale?’. A hundred well-chosen interventions deployed independently produce a zoo of pilots: inconsistent data definitions, duplicated agents, no shared supervision layer, and no way to evidence to a regulator how AI participated in an engagement.
This is where the AI-native operating system we described in our earlier paper does its work. The shared domain ontology gives agents and auditors a common language for accounts, assertions, risks, and evidence. The agent marketplace distinguishes reusable capability agents from audit-domain agents and engagement-specific niche agents — which, usefully, is also an inventory that makes agentic washing visible: an ‘agent’ that cannot be described in terms of its goals, tools, and guardrails is more likely a well-built workflow — valuable, but governed differently. Model context protocol servers govern how agents reach client systems and external data. And agentic workflows compose individual task-level interventions into end-to-end audit procedures, with the autonomy level of every step recorded — which is precisely the documentation an inspection will one day ask for.
The autonomy model and the operating system are two halves of one design discipline: the model decides what each task needs; the operating system makes those decisions consistent, governable, and reusable across the practice.
How do organisations determine which audit tasks are best suited for AI agents, assistants, or traditional automation?
For audit leaders, the implication is a programme that begins with neither a vendor selection nor a proof of concept, but with an inventory.
Decompose the methodology into tasks. Place each task on the autonomy model, recording the rationale. Be candid about the tasks that need no agent at all. Concentrate early investment where Level 3 expands coverage and quality — journal entries, evidence scanning, estimates support. Build the supervision and documentation patterns now that Levels 4 and 5 will eventually require. And ask of every ‘agentic’ claim — internal or external — the questions the profession asks of any assertion: what does it actually do, and what is the evidence?
Our experience is that the teams who decompose the work first tend to adopt more confidently and govern more comfortably — and they are rarely caught out when clients and regulators ask the question that is surely coming: where, exactly, did AI act in this audit, and who was watching?