Responder
defaultGrants nothing. Answers on request; writes only on direct order, confirmed in the same exchange.
How autonomous an AI agent is, stated in a unit you can enforce in code and check as a customer.
No jargon: what problem the scale solves, what each level actually grants, and why the ceiling is set by the risk of the action rather than by how clever the agent looks. Recorded in Spanish.
The word «autonomous» carries no measurable content in enterprise software. A product that drafts an email and a product that files a tax document without review are both marketed as agentic. The buyer has no unit in which to ask «how autonomous?», the operator has no dial to set, and the auditor has no artefact to inspect.
Ordinal scales of automation have existed since 1978, and one reached the market: the six driving-automation levels of SAE J3016, which succeeded by allocating responsibility, not by ranking intelligence. For AI agents there are two recent antecedents —Feng, McDonald and Zhang (2025) and the Cloud Security Alliance (2026)— both earlier than this scale, and both remain descriptive: they classify a deployment after the fact. Neither specifies how a level becomes a runtime constraint, who is entitled to change it, or what the customer is owed when the agent acts alone.
That is what this scale contributes: the operational construction, not the idea that autonomy admits levels.
They come from composing two axes that vary independently: initiative (who starts the loop) and execution authority (what the agent may commit without a human). They are ordered by the smallest increment of authority granted, not by how sophisticated the behaviour feels.
Grants nothing. Answers on request; writes only on direct order, confirmed in the same exchange.
The right to address you unprompted: scheduled reports, event-anchored check-ins. Still executes nothing.
The right to prepare work: concrete actions queued for a human to approve or dismiss one by one.
The right to commit low-risk actions unaided, audited and reversible within the stated window. Without those two guarantees, this level does not exist.
The right to sequence those commits towards an assigned objective, under run caps, blocking escalation and a kill switch. It grants sequencing, not new commit authority.
The right to delegate to other agents under a coordinator, inheriting every ceiling above. Specified, not validated.
The right to act across organisations: one company's agent coordinating with another's. Not more of the same —it moves a third axis, the authority domain— and it is governed by composition rules: authority = intersection of both catalogues, ceiling = union of both ceilings, undo = the shorter of the two, bilateral audit, unilateral descent.
V5 and V6 are specified and reserved: their semantics are defined but not yet exercised in production, so nobody can claim conformance to either today. They are specified for the reason SAE J3016 defined level 5 years before any vehicle could claim it: a scale is normative, and the level that does not yet exist is the one that tells implementers what they are building towards.
Only two transitions deserve their own price and audit. V2 → V3 is the only one that changes who commits a write; everything earlier changes only who speaks first. V3 → V4 does not remove the human decision, it removes its point: the human sets the objective and the agent decides the sequence.
In existing frameworks the level governs the whole agent. Here the decision is a matrix of level × the action's risk class, under a ceiling no level overrides. The unit of risk is the action type, not the request, and it is assigned by four criteria: technical reversibility, externality, legal or monetary effect, and third-party visibility.
Internal, fully reversible, no accounting or legal effect.
create task · update task · internal message · toggle checklist item · save learned note · create reminder · schedule meeting
Internal but consequential to shared state: duplicates, reports, other people's work. Reversible with effort.
create contact · create event · set attendance · change vendor status · create email draft
Money, fiscally or legally effective documents, and any communication leaving the organisation. Irreversible in the sense that matters: it has already been read.
fiscal document · register payment · reconcile transaction · dunning email · send or reply to email · reply to lead
if the action is in the invariant ceiling → human (at any level, V5 included) if the risk is high → human if the risk is low and level ≥ 3 → execute, audited and undoable otherwise → propose for approval action absent from the catalogue → treated as high risk execution that throws → degrades to a proposal, never to silence
In the Hubents catalogue, 10 of 29 write actions (34 %) are high risk and 13 (45 %) are low. Under a single dial, an operator who wants internal task upkeep automated must either accept client-facing email too, or forgo automation entirely. Under the matrix, the 45 % that is internal and reversible is exactly what V3 grants, and the 34 % that is external or financial stays behind a human at every level.
This is what makes it defensible to sell autonomy in levels rather than «more intelligence»: what the customer buys at each level is a bounded, enumerable set of commitments.
What separates an operational scale from a marketing label. A system is Vn-conformant if it satisfies all seven and publishes its statement.
Vn ⊂ Vn+1. No level trades away an oversight affordance for authority.
A non-empty, published set of actions requires a human at every level.
Every autonomous action produces a record and, where reversible, an undo affordance within a stated window.
Initial level V0. An uncatalogued action is maximum risk. A failure degrades to a pending proposal, never to silence.
Lowering takes effect at once. It is the kill switch: it must never be harder than ascent.
The level is set by the customer, in the language of the business, and reserved levels are visibly marked.
Gates and risk classes live in program logic, under regression test. Not in the prompt.
An agent instructed «never send invoices without confirmation» holds that boundary at the pleasure of the model. The same boundary written as a branch in the execution path holds regardless: whatever the model decides, whatever a user writes into a prompt, whatever an injected document requests. A prompt is a request; a gate is a constraint.
A level claim is worth as much as its verifiability. Ten fields the vendor
publishes and a customer or auditor can check against the running system.
There is a template,
a JSON Schema, and a
single-file, zero-dependency reference implementation:
vaas.mjs — the gate, the validator
and the matrix, ready with node vaas.mjs validate.
It has two virtues: it is falsifiable —a customer attempts a high-risk action at the declared level and observes whether a human is required— and it is comparable: two vendors claiming «we are V3» can be read side by side, and the difference shows up in fields 5, 6 and 7 rather than in adjectives.
The scale is not a proposal on paper: it comes from two multi-tenant SaaS products that needed a dial. They differ in every free parameter and agree in every fixed one, which is the intended shape of two conformant systems.
| Hubents · SofIA | Koble · KobIA | |
|---|---|---|
| Domain | Event management | Agency operations / CRM |
| Levels offered | V0–V4 | V0–V4 |
| Reserved | V5 · V6 | V5 · V6 |
| Self-service up to | V4 | V3 |
| Actions classified | 29 · 13 / 6 / 10 | 21 · 11 / 5 / 5 |
| Autonomous-execution threshold | V3 | V3 |
| Undo window | 7 days | 7 days |
| Invariant ceiling | 10 actions | 5 actions |
| Agent in production since | 17 Feb 2026 | Jul 2026 |
| Scale added | Aug 2026 | Aug 2026 |
| Statement published | /autonomia | /autonomia |
nobody ends up autonomous by accident
across both products, none reversed
the queue is the weak point, not the executor
a granted level is not an exercised one
Hubents · SofIA. Read from the database, not from recollection: 196 active organisations (193 at V0, 1 at V1, 1 at V2, 1 at V4), agent in production since 17 February 2026 with 270 conversations, and 270 proposals between 14 June and 11 August: 161 dismissed (60 %) with no reason recorded, 95 pending (35 %) and 14 autonomously executed, none reversed.
Active means active: 14 deleted accounts and 1 suspended one are excluded. The per-level counts in both statements use the same criterion, so the numbers on this page and the ones the products publish can be cross-checked.
Koble · KobIA. Three active workspaces —2 at V0 and 1 at V3— and 20 proposals in its first week: 19 executed, 17 of them without a human, 1 pending, none dismissed and none reversed. It is the mirror image of Hubents, and that is why we publish it: same autonomous-execution threshold, same undo window, opposite queue behaviour. With a single workspace at V3 there is nothing to conclude yet; we state the number, not a trend.
A granted level is not an exercised one: the V4 organisation has been there for months and has run zero missions; its actual autonomy is V3's. No framework that classifies by configured level sees that difference.
And the weak point is the queue, not the executor: while the 31 autonomous commits across both products produced no reversals, Hubents' proposal queue accumulates and gets dismissed. That is automation bias arriving from the direction we did not expect — not humans rubber-stamping proposals, but humans abandoning them. We publish it prominently because it is the number least flattering to the scale.
Under CC BY 4.0: use it, translate it, adapt it to your domain and criticise it, commercially included. Attribution is the only requirement.
There is no certification body and none is needed: conformance is self-declared and falsifiable. A false claim is refuted by its own document. Full conditions in the name policy.
The United Nations does not legislate on AI, and this scale claims nothing of the sort. What exists is a mechanism for shared assessment: Resolution A/RES/79/325 (26 August 2025) established the Independent International Scientific Panel on AI —forty members appointed in February 2026— and the Global Dialogue on AI Governance, whose first session met in Geneva on 6–7 July 2026. The Panel's remit is evidence-based scientific assessment, reported annually.
That remit has a measurement problem, and it is the one this page opens with: an assessment of how autonomously agents are deployed cannot be assembled from vendor self-description, because «autonomous» is not a comparable quantity across products or jurisdictions. What such a process can consume is a unit: a level with fixed meaning, an enumerable grant, and a claim falsifiable by inspection rather than by trust.
The honest limit is the usual one: this is one instrument, validated in two products of a single author and published for criticism. If forty scientists find a better unit, the useful outcome is that a unit exists to argue about.
The scale is named after its author for lack of an existing standard to extend. The genealogy —Sheridan and Verplank 1978, SAE J3016, Feng et al. 2025, Cloud Security Alliance 2026— is acknowledged deliberately: what is claimed is the operational, risk-gated construction, not the idea that autonomy admits levels.
No mile-long forms. No silly questions. Just enough to know if we fit and get back to you fast.
I reply personally within 48 hours. If you don't hear from me in that time, something broke.