Articles /

Part 1: What Developers Think They Use vs. What They Actually Use: Measuring AI Honestly

Developers flag only about half the AI work they actually do. Our February 2026 productivity data shows where self-reported adoption numbers break down, and why delivery outcomes improve even on tasks nobody ever tagged as AI-assisted.

a blue and white cover with a robot hand

Developers feel they’re moving faster as they use AI, so we looked at the data again. But as our data told us, that feeling didn’t show up supporting it. That’s only one of the many unexpected yet exciting findings we had back in our November 2025 AI productivity findings.  

A few months passed and we ran the same measurement again. We had more data this time to look at what really happens when our developers use AI in their work and how AI improves that same work. 

This is the first of three blogs' series on what we found in the AI productivity report findings for February. In part 1 of this series, we discuss measuring AI honestly, because every finding that follows depends on getting the measurement right. Part 2 shows why adoption is a capability you build, not a tool you switch on. Part 3 covers why fixed-price delivery gets safer with AI.

Developers were reporting only about half of the AI work they were actually doing. We found this out when we stopped depending on what people told us and started measuring AI practice directly. If you want to grasp AI's impact on your delivery, you have to figure out the measurements first before heading into analysis.

Why self-reporting wasn't enough 

Our first round of measurement relied on a simple mechanism where developers ticked a box on their timesheet when a task was AI-assisted. It was a sensible start as it gave us a usable signal. But a self-reported flag only records what a person remembers to note down, and only what they personally count as "AI work." 

So, in this round we implemented another independent source, ground-truth usage data. This data was captured automatically at the session level, without anyone needing to remember anything. That gave us something to check the self-reports against with a way to not just ask "what do developers say they're doing with AI" but "what are they actually doing." 

The two pictures did not match. 

Developers under-report their AI use by roughly half 

From the cross-referencing automatic usage data against the self-reported flags, it revealed a uniform and systematic gap. A large portion of the sessions where AI was demonstrably in use carried no AI flag on the parallel timesheet entry. In round numbers, developers were flagging only around half of their actual AI usage. This wasn't forgetfulness or dishonesty, it was systematic, and the pattern told us why. 

Only 36% of observed AI activity was flagged as AI-assisted. A further 27% happened during timesheeted work without being reported, and 36% fell outside any logged time.

How much AI use actually gets recorded; only about half of real AI sessions are logged as AI-assisted.

The sessions people forget to flag are the light ones 

When we looked at what differentiated a flagged session from an unflagged session, a clear example appeared. Flagged sessions were the heavy, deliberate ones which loads a large amount of context and generates substantial output, this was clearly "AI work." Those then get recognised and recorded. 

The unflagged sessions were lighter and more conversational. You would ask a question, get an answer, clarify a problem, and move on. That kind of use no longer feels like AI assistance, it feels like working. And it's precisely because it has become normal, and it goes unrecorded. 

This is an important signal, when a tool stops feeling like a tool and starts feeling like part of how you think, people stop noticing they're using it. The under-reporting is itself evidence of how deeply embedded AI has become in everyday delivery work. 

The implication is that usage stats from self-reporting are directional, not absolute 

Two things follow, and both matter for anyone trying to measure AI in their own organisations. 

Firstly, any adoption figure built on self-reported flags is an undercount, and a biased one, skewed toward heavy, visible sessions and blind to the constant low-level assistance that now runs underneath everything. If your AI adoption dashboard says 20%, the real figure is materially higher. 

Second, and more interesting, the benefits of AI show up even on work that was never flagged as AI-assisted. When we compared task outcomes from before broad AI adoption against the period after, the distribution of outcomes shifted even across tasks that carried no AI flag at all. This shift was tighter with fewer tasks running catastrophically long. 

We call this the ambient effect. AI lifts the floor on everything, like faster lookups, less context-switching, and lower friction on the small sub-problems that derail an estimate. Measure only the explicitly tagged work and you will systematically understate AI's true impact. 

Before and after broad AI adoption. Left: outcome distribution tightens sharply around on budget. Right: overruns drop across all four severity bands, from any overrun through to severe.

Here’s how task outcome distribution look before and after broad AI adoption. The after-adoption distribution is tighter with a shorter tail, even across tasks not flagged as AI-assisted.

Why this matters before any analysis 

There is a strong temptation, when leadership asks "is AI working for us?", to answer with a survey or a self-reported "hours saved" number. We did that, and we learned the hard way that it produces a confident answer to the wrong question. People are poor at estimating a counterfactual, like how long something would have taken, and they are poor at noticing assistance that has become routine. 

The honest path is harder but far more defensible here. We measure what actually happened in delivery, from data that doesn't depend on memory or self-assessment, and we let the outcomes speak. That's the foundation everything else in this series is built on. 

AI Measurement: Key Findings 

Finding 1: Self-reported AI usage undercounts reality by roughly half. When checked against automatic, session-level data, developers had flagged only about half of their actual AI-assisted work. 

Finding 2: The gap is systematic, not random. Heavy, deliberate sessions get recognised and recorded. Light, conversational use has become invisible because it now feels like ordinary work. 

Finding 3: Under-reporting is itself a maturity signal. When a tool stops feeling like a tool, people stop noticing they use it. The blind spot is evidence of how embedded AI has become.   

Finding 4: AI's benefit appears even on un-flagged work. Task outcomes improved across the board after broad adoption, instead of just explicitly AI-tagged tasks. Measuring tagged work alone understates the real effect. 

Finding 5: Measurement has to come before analysis. Surveys and self-reported savings answer a question people can't answer accurately. Ground-truth delivery data is the only defensible starting point. 

AI Measurement FAQs 

Q: Why can't we just ask developers how much they use AI?  

Because self-reports drift low in a predictable way. Your team would need to switch from asking developers to capturing usage directly from your work environment. Routine AI assistance stops feeling like "using AI" and goes unrecorded, while time-saving estimates get inflated. Ground-truth measurement removes both biases and gives you numbers you can actually act on. 

Q: What does "ground-truth" usage data mean?  

It’s the instrument in your delivery environment to capture AI usage automatically, at the point of work. When usage is logged by the system rather than recalled by the person, you get a record that reflects what actually happened. It’s not what someone remembered or thought to report, and that's the data worth building decisions on. 

Q: If usage is under-reported, are our adoption numbers wrong?  

Treat your current adoption figures as a floor, instead of a ceiling, and then close the gap. The trend and relative differences between teams are still useful, but the absolute number is higher than self-reporting shows, often substantially. Use that as the starting point for a more accurate measurement approach, and you'll find adoption is further along than you thought. 

Q: What is the "ambient effect"? 

It's the improvement in delivery outcomes that shows up across all work once AI is broadly available. You can design your measurement to capture all work, and not just the tasks explicitly flagged as AI-assisted. Once AI is broadly available in a team, it lifts delivery outcomes across the board, including on tasks where no one pressed a button to record it. Widen the lens, and you'll see the full picture. 

Q: Where should an organisation start if it wants to measure this properly?  

Start with your own delivery data and build a measurement model around what actually happened, not what people reported. Pull your historical task-level data, establish a baseline, and add automated usage capture going forward. These will give you a before-and-after you can defend, and a foundation for every AI investment decision that follows.

Measuring AI Honestly in Your Organisation 

Most organisations are measuring AI adoption the way we started and they’re quietly drawing conclusions from numbers that don't hold up. And this happens even with self-reported flags and "hours saved" estimates. 

We've been through that and, with these findings, we’re confident we can conduct the same study and analysis for your organisation with your own, reliable data. 

And with the amount of data that we gathered and analysed, we’ll be sharing more findings in the weeks to come. You’ll find out how AI adoption really works from a before-and-after comparison. You’ll also learn how AI makes fixed-priced delivery safer. So stay tuned for our next few blogs to learn more insights. 

Measure AI Use Your Way

If you’re eager to learn our findings right now, our team of AI experts is just one discovery call away. Schedule a free call and learn how AI can create value for your team.

a group of people sitting at a table with laptops

Elevating Excellence: Our Microsoft Partnership!

Our ongoing collaboration with Microsoft brings you the best in innovation and technology.

Tell us about your project

Got a question, need support, or ready to explore possibilities? We're here to help. Fill out the form and our team will reach out within 1-2 business days.