6 data sources · 24 measurements · Updated monthly

AiR Bench

Are you ready for the Ai Revolution?

Compare your business against the latest frontier understanding of AI’s impact, and see exactly how to optimise for it.

Grounded in frontier research from

MITMcKinsey & CompanyCiscoStanfordAnthropicDORA

The case

AI rewrites the rules every month. A one-off survey cannot keep score. A living benchmark can.

AiR Bench measures what your organisation does with AI, not how it feels about it. We re-score you every month as the frontier moves, so your standing always stays current.

01

Outcomes, not opinions

We ask behavioural questions with observable anchors. The benchmarks that last, like DORA, measure what teams do, not how confident they feel.

02

The benchmark is alive

It is versioned like software. Every score is timestamped to an instrument version. When the frontier moves, your standing can drift even if nothing you do changes. We show you that drift.

03

Show your workings

We publish every question with its rationale and evidence. Weights, thresholds, the changelog, the conflicts of interest: all public. We earn trust by being more open about method than anyone else.

One instrument · Two axes · 24 scored measurements

Six readiness dimensions. Then the block nobody else measures.

D1

Strategic Clarity

Whether AI ambition has been converted into named, owned, quantified opportunities.

D2

People and Change

Leadership engagement, frontline involvement, and actual weekly use across the organisation.

D3

Process Readiness

Whether the work AI will touch is understood, documented and being redesigned, not just augmented.

D4

Data Foundations

Whether the data the use cases need is accessible, usable and governed.

D5

Governance and Trust

Whether use is governed by policy people actually know, with defined human review and explainability.

D6

Execution Track Record

Whether the organisation ships digital change, and how fast.

R · THE REALISATION AXIS

Value Realisation

What no competitor measures

We look at production deployments, pilot-to-production conversion, measured value against baselines, where that value concentrates, and how fast the frontier reaches your workflows. Current research keeps finding the same gap between adoption and value. An instrument that cannot see that gap cannot claim accuracy.

The Curve · Updated monthly

Adoption is sprinting. Value is walking.

See how AI adoption converts to value over time. Three lines track adoption, production and measured value, annotated with model releases, research findings and instrument versions. It is the visible heartbeat of the research layer.

Explore The Curve →
02550751002023202420252026ChatGPT momentAgentic wave beginsGenAI Divide publishedAiR v1.0
Using AI somewhereProduction use casesMeasured value% of organisations · Sourced from the Evidence Register · Updated monthly

The living research layer · The differentiator

Most assessments are snapshots. This one has a heartbeat.

Every month the research loop runs. New studies, model releases and market evidence enter the Evidence Register, and a standing review decides what counts as signal. Signal becomes the next instrument version. When the instrument re-anchors, we re-score your stored answers under it.

Your answers stay the same and the frontier moves. That gap is drift, and this is the only instrument that shows it to you.

Last review

2 June 2026

Next review

6 July 2026

v1.1 expected

Q4 2026

From the Evidence Register · Live feed

12 Jun

Moonshot AI Kimi K2.7 Code and continued open near-frontier coding releases (release-tracker aggregation)

watching

Moonshot's Kimi K2.7 Code extends the run of low-cost, open-weight, near-frontier coding/agentic models (alongside earlier DeepSeek V4), lowering the economic barrier to agentic build-out for cost-sensitive mid-market firms; sourced from release trackers rather than a single primary lab post, so treated as a developing signal. Informs R5 (frontier responsiveness) and R6 (agentic deployment cost).

08 Jun

Microsoft AI MAI model family and 'Frontier Tuning' (Microsoft Build 2026 / Microsoft AI announcement)

watching

Microsoft AI announced seven in-house multimodal models (incl. MAI-Thinking-1, ~35B active MoE, 256K context, 53% SWE-Bench Pro) and 'Frontier Tuning' — reinforcement learning on a customer's own workflow traces inside their environment. Signals a shift toward customer-data-as-moat and tunable near-frontier models; informs R5 (frontier responsiveness) and D4 (Data Foundations as the input to tuned models). Primary vendor source but partly product marketing, so monitoring independent benchmarks.

02 Jun

White House Executive Order,'Promoting Advanced Artificial Intelligence Innovation and Security' ((U.S. federal government)

entered

Establishes a voluntary framework for federal early access (up to 30 days) to 'covered frontier models' and a classified NSA/CISA/Treasury benchmarking process for cyber capability, explicitly disclaiming any mandatory licensing or preclearance; deliverables due 1 Aug 2026 (corroborated by Wiley law alert and multiple outlets). Sets the near-term US governance posture against which UK mid-market buyers judge vendor risk. Informs D5 (Governance & Trust) and R5 (frontier responsiveness).

02 Jun

Anthropic Economic Index, May 2026 update

entered

Reliability-adjusted productivity contribution revised upward for integrated deployments. Informs the R5 frontier-responsiveness anchors for v1.1.

02 Jun

EU AI Act implementation guidance, second tranche

watching

Watching. Explainability obligations may re-anchor D5.3 — what counts as 'could explain to a regulator, today' is about to get a legal definition.

01 Jun

GitHub Copilot move to usage-based billing ('AI Credits') (GitHub/Microsoft)

watching

Copilot shifted from flat per-seat pricing to token/credit metering across all tiers, with heavy agentic users reportedly paying $60–100+/month versus the $10 sticker (verified against GitHub's pricing page in trade coverage). Concrete evidence that agentic deployment carries variable, usage-scaled cost that reshapes pilot-to-production economics. Informs R6 (agentic deployment), R2 (pilot-to-production conversion) and D3 (Process Readiness).

Monthly research review. Minor versions roughly twice a year, driven by the register, never by the calendar. · Every change published in the changelog

Two ways in

AiR Bench

10 to 12 minutes · No cost

Benchmark your business against the frontier. You get your band, the AiR Grid, seven dimension percentiles, the perception gap, and a clear Now, Next, Later.

Benchmark your business →

AiR Advisory

The benchmark plus one session

We check your answers against real artefacts with our consultants. You get a badge valid for 12 months, and a separate stratum in the public dataset.

Talk to us →

The covenant

Your results are private. Always.

We never publish, share or sell an individual or identifiable organisational result. We do publish aggregated, anonymised data as a public resource: sector cuts, band distributions, the perception gap, and the readiness-to-realisation relationship.

You can delete your data at any time. Deletion carries through to the aggregates at the next publication cycle.

Part of the product, not the small print. Read it in full.

6 data sources · 24 measurements · Updated monthly

AiR Bench

Twelve minutes now. A score you can defend in the boardroom.

Benchmark your business →

See your headline before any sign-up · Save and resume · v1.0.0