Engineering Metrics for CTOs: What to Report Upward
By the CTO Coach TeamReviewed against primary sourcesPublished
A CTO should report a small set of delivery and reliability trends to the board, a wider set to the executive team, and the full diagnostic detail only to the engineering teams that own it. DORA's five software delivery metrics are the best-evidenced starting point, and the SPACE framework explains why no single number is enough.
This guide explains what the five DORA metrics measure, which audience gets which metric, which measures backfire, and how to measure the effect of AI tools. It draws on DORA's own guidance and the SPACE paper's published framework, and it labels our editorial judgement where the sources stop. For a ready-made board format, use the CTO board update template.
What are the DORA metrics in plain English?
DORA, the Google Cloud research programme, defines five software delivery metrics in two groups. Throughput covers change lead time, deployment frequency and failed deployment recovery time. Instability covers change fail rate and deployment rework rate. DORA's research finds that speed and stability are not trade-offs: for most teams, the two groups move together.
| Metric | DORA definition | Group | Plain-English version for non-engineers |
|---|---|---|---|
| Change lead time | Time for a change to go from committed to version control to deployed in production | Throughput | How long a finished idea takes to reach customers |
| Deployment frequency | Number of deployments over a period, or time between deployments | Throughput | How often we ship |
| Failed deployment recovery time | Time to recover from a deployment that fails and needs immediate intervention | Throughput | How fast we fix a release that goes wrong |
| Change fail rate | Ratio of deployments that need immediate intervention afterwards | Instability | How often a release causes trouble |
| Deployment rework rate | Ratio of deployments that are unplanned and result from a production incident | Instability | How much of our shipping is cleaning up after incidents |
Source: DORA, software delivery performance metrics guide. DORA's current set has five metrics; older articles describe four "key" metrics and use mean time to restore, which the current model replaced with failed deployment recovery time. If a vendor page still lists four, check its date.
DORA also publishes usage guidance that matters more than the definitions. It says to apply the metrics to one application or service at a time, because blending them across teams can mislead; to interpret them in your own context rather than compare teams; not to set them as targets, because mandated numbers invite gaming (the guide links this to Goodhart's law); and to use several together, including some that create healthy tension. It recommends starting with conversations or the DORA Quick Check before investing in complex integrations.
Which metrics should go to the board, the executive team and the engineering team?
Show the board two or three outcome trends, show the executive team the delivery system and its risks, and keep team-level diagnostics with the teams. The rule is that each audience should see only the measures it can act on. The table is our editorial recommendation, built on the DORA and SPACE guidance cited in this guide, not a published standard.
| Board | Executive team (CEO, CFO, product, sales) | Engineering team | |
|---|---|---|---|
| Question they ask | Is the plan on track and is the company safe? | Can we commit to dates, and what is slowing us down? | What do we fix next? |
| Delivery metrics | Two or three of the DORA five as a quarterly trend, for example change lead time and change fail rate | All five DORA metrics per major product, as a trend | All five per service, reviewed in retrospectives |
| Reliability and risk | Customer-impacting incidents, with a one-line cause and fix | Incident trend, recovery time, top open risks with owners | Incident reviews, alert quality, on-call load |
| Delivery against plan | Share of committed outcomes delivered, red/amber/green | Roadmap progress and slippage by cause | Sprint or cycle commitments and carry-over |
| People | Headcount against plan, key attrition | Hiring pipeline, retention, engagement signals | Team health surveys, workload |
| Money | Engineering spend against budget | Spend by product or platform, cloud cost trend | Cost per service where relevant |
| Quality of the system | Technical debt as a business risk, if material | Debt work as a share of capacity | Debt backlog, code review time, build health |
Two practical rules follow. First, report a trend, not a snapshot: one deployment-frequency number is meaningless without the previous four quarters. Second, translate every metric into a business consequence, as covered in board-ready CTO communication. The template linked above structures the same choices in one page.
What does the SPACE framework add to DORA?
DORA measures the software delivery pipeline. SPACE is broader: it argues that developer productivity cannot be captured by a single metric or dimension. The framework, published in ACM Queue in 2021 by Nicole Forsgren, Margaret-Anne Storey and colleagues at GitHub and Microsoft Research, names five dimensions.
| SPACE dimension | What it covers | Example signal (ours, illustrative) |
|---|---|---|
| Satisfaction and well-being | How developers feel about their work, team, tools and culture | Quarterly developer survey |
| Performance | Outcomes of a system or process | Customer-visible reliability, features adopted |
| Activity | Count of actions or outputs | Deployments, merged changes |
| Communication and collaboration | How people and teams work together | Review turnaround, onboarding time |
| Efficiency and flow | Ability to make progress with minimal interruption or delay | Time lost to waiting, interruptions |
Sources: the paper's abstract on Microsoft Research states that productivity "cannot be measured by a single metric or dimension"; the dimension definitions are summarised in a DX primer by Abi Noda. DX sells developer productivity products, so we used it only for the dimension names and definitions, and we recommend reading the original paper for the full guidance. The primer also notes that the paper's example metrics are examples, not recommendations, and that activity alone is insufficient because software development is knowledge work.
The useful lesson for a CTO is to pair a delivery metric with a human one. A team whose deployment frequency rises while its survey results fall is telling you something the DORA numbers alone would hide.
Benchmark the budget too: the free engineering budget benchmark compares your R&D spend as a share of ARR with the SaaS Capital median and shows the formula.
Which engineering metrics backfire?
Metrics backfire when they are used as targets or applied to individuals. DORA warns that mandating numbers invites gaming, and the SPACE authors warn against measuring activity alone. Lines of code, story points per person and commit counts are classic examples, and our editorial view is that they should stay out of any report that leaves the team.
| Metric | Why it misleads | Better substitute |
|---|---|---|
| Lines of code or commit counts | Activity, not outcome; rewards volume over simplicity | Change lead time, customer-visible outcomes |
| Story points per person | Points are a team estimating tool; comparing people or teams makes estimates inflate | Delivery against committed outcomes |
| Individual deployment counts | Delivery is a team and system property | Team-level DORA trends |
| Utilisation or "hours worked" | Rewards busyness; hides waiting and rework | Efficiency and flow signals from the team |
| A single "productivity score" | SPACE says one metric cannot capture productivity | A small set across several dimensions |
| Team league tables | DORA says results are context-specific and not directly comparable | Each team against its own trend |
Technical debt is a separate case, because it is hard to metricise but expensive to ignore. In the 2024 Stack Overflow Developer Survey, 62.4% of professional developers named technical debt as their top frustration at work (Stack Overflow, 2024). Report debt to the board as a risk with a cost, and track it with the team through its effect on lead time and rework. The first tech debt crisis guide covers the conversation, and the tech debt calculator helps rank items.
How do you measure the impact of AI tools on engineering?
Record a baseline for the DORA metrics and a developer survey before rollout, then compare after one and two quarters. Do not rely on adoption numbers or self-reported speed-ups. The public evidence is mixed, and the DORA data suggests AI amplifies existing strengths and weaknesses.
DORA's 2024 report associated higher AI adoption with an estimated 1.5% lower delivery throughput and 7.2% lower delivery stability, while a 25% increase in AI adoption was associated with a 3.4% improvement in code quality (Google Cloud, 2024). The 2025 report found 90% adoption but only 24% of respondents trusting AI a lot or a great deal (Google, 2025). Both are surveys, so they show association, not cause. The METR randomised trial is the cautionary example for self-reports: experienced developers took 19% longer with AI tools after predicting a 24% speed-up (METR, 2025), though METR now marks it out of date.
Our AI coding tools evidence review covers each study's design and sample. For the board-level framing, see CTO AI strategy.
How do you start a metrics programme without a big tooling project?
Start with one product and the DORA Quick Check, then add measures as questions arise. DORA's own advice is not to let measurement crowd out improvement. A sequence we recommend, as editorial guidance:
- Pick one service or product and agree with its team what each of the five metrics means there. Definitions vary, for example what counts as a deployment.
- Capture a baseline from the data you already have: the deployment pipeline, the incident tracker and the version control history.
- Add one human measure, such as a short quarterly developer survey, so the numbers have context.
- Share the results with the team first, and discuss what to improve, before any number leaves engineering.
- Choose the board pair (for example change lead time and change fail rate) and show them as a trend with one sentence of meaning.
- Review quarterly and retire any metric nobody acts on.
This is also the point where a new CTO earns credibility with a non-technical CEO. The first 90 days as CTO guide includes a week-by-week plan for establishing the baseline, and scaling an engineering org explains how the metrics you need change as the organisation grows.
Frequently asked questions
What are the DORA metrics?
DORA's current model has five software delivery metrics: change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate. The first three measure throughput and the last two measure instability. DORA says to apply them per service, avoid using them as targets and not to compare teams.
Which engineering metrics should a CTO show the board?
Two or three outcome trends, not a dashboard. A common pair is change lead time and change fail rate, shown over several quarters with one sentence of business meaning, plus customer-impacting incidents and delivery against plan. Keep team-level and individual measures out of the board pack and in an appendix.
What is the SPACE framework?
SPACE is a developer productivity framework published in ACM Queue in 2021 by Nicole Forsgren, Margaret-Anne Storey and colleagues. It names five dimensions: satisfaction and well-being, performance, activity, communication and collaboration, and efficiency and flow. Its central point is that productivity cannot be measured by a single metric.
Are story points or lines of code good metrics?
No, not as performance measures. Both count activity rather than outcomes, and the SPACE authors warn against measuring activity alone. DORA also cautions that numeric targets invite gaming. Use story points as a team planning aid and measure outcomes through delivery metrics and customer impact instead.
How should I measure the impact of AI coding tools?
Set a baseline for the DORA metrics and a developer survey before rollout, then compare after one and two quarters. Do not rely on self-reported speed-ups: METR's 2025 trial found developers took 19% longer with AI tools despite predicting a 24% speed-up. DORA's reports show association, not causation.
Sources
- DORA metrics: the four keys (current five-metric model) DORA, 2026
- DORA research programme DORA (Google Cloud), 2026
- The SPACE of developer productivity: there's more to it than you think Microsoft Research (ACM Queue, 2021), 2021
- SPACE framework: a quick primer (dimension summary; vendor source) DX, 2026
- Developer Survey 2024: professional developers Stack Overflow, 2024
- Announcing the 2024 DORA report Google Cloud, 2024
- DORA report 2025: State of AI-assisted Software Development Google, 2025
- Measuring the impact of early-2025 AI on experienced open-source developer productivity METR, 2025
How we source and check figures: Methodology.
Ready to level up?
Discover your strengths and gaps with our free CTO Readiness Assessment.
Take the CTO Readiness Assessment