Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

Are “Ghost Engineers” Stunting Software Productivity? What the 9.5% Claim Really Shows

A Stanford-linked effort reported a 9.5% “ghost engineer” estimate, but its method and interpretation remain contested. The figure is not a settled industry rate or proof that flagged engineers do no useful work.
From TheFinanceBase Team5 min to read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Not conclusively. A Stanford-linked research effort estimated that 9.5% of software engineers in its dataset produced less than one tenth of the median engineer’s measured output. That is a striking early finding, not a settled estimate of how many engineers do little work or a basis for calling them a drain on their employers. The model looks at code changes, and important engineering contributions may not appear in them.

What does the “ghost engineer” claim mean?

In 2024, researcher Yegor Denisov-Blanch described engineers in the lowest-output category as “0.1x-ers,” saying they did “virtually nothing.” ITPro reported the estimate as 9.5% of software engineers and said the dataset covered more than 50,000 engineers at hundreds of companies. The claim concerns measured output relative to the median in that dataset; it does not establish that those engineers performed no work.

The distinction matters. A relative score from one dataset is not automatically a representative estimate for the software workforce as a whole. The figure should be read as a result reported by one research effort, rather than a universal industry statistic or a verified count of employees who contribute nothing.

How does the productivity model work?

The Stanford Software Engineering Productivity Research site describes an approach that uses machine learning to approximate expert evaluations of software commits. Denisov-Blanch has said the system analyzes source-code changes in private Git repositories and simulates a panel of 10 expert reviewers. He has also clarified that it is intended to assess the substance and complexity of changes, not simply count commits or lines of code.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The related 2024 arXiv preprint, Predicting Expert Evaluations in Software Code Reviews, reports correlations of r = 0.82 for coding time and r = 0.86 for implementation time between model predictions and expert evaluations. Those figures describe agreement on particular review dimensions; they do not show that the model accurately identifies every low-performing employee or establishes the 9.5% estimate as a population-level fact. The abstract says the work addresses review dimensions “typically avoided due to their complexity or subjectivity.”

What did the study report about remote work?

Denisov-Blanch’s 2024 public thread reported different shares of engineers in the “ghost” category by work location:

Work arrangement Reported share How to interpret it
Fully remote 14% Author-reported subgroup estimate from the same research effort; not an independently established prevalence rate.
Hybrid 9% Author-reported subgroup estimate from the same research effort; not an independently established prevalence rate.
Office-based 6% Author-reported subgroup estimate from the same research effort; not an independently established prevalence rate.

The figures suggest a difference within the reported dataset, but they do not show that remote work causes lower productivity. The public account does not establish that the groups were comparable in role, seniority, company, project, or other conditions. The percentages therefore should not be used to rank work arrangements or predict an individual engineer’s performance.

Why is the 9.5% figure contested?

Jellyfish noted that the viral estimate had not been peer reviewed and that neither the paper nor the researcher’s post clearly explained how the 9.5% figure was derived. Pluralsight raised questions about how a larger dataset—1.73 million commits and 50,935 engineers—related to model training. It also argued that commit-centered analysis can overlook mentoring, debugging, architecture, code review, and other work.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These critiques do not prove the estimate wrong. They identify limits on what can be concluded from the public account: the method behind the specific threshold and estimate is not described clearly enough there to treat the result as independently validated workforce prevalence. The dataset figures reported by ITPro and discussed in the critique also should not be assumed to describe the same analytical stage or sample without a clear explanation of that relationship.

Denisov-Blanch has said participating organizations checked many flagged engineers and that ancillary activities did not explain most cases. That is the author’s account, not independent confirmation of the estimate. He has also cautioned that “Decisions shouldn’t be made purely based on what our model spits out.”

What can different productivity measures actually tell an employer?

No single metric captures every part of engineering work. Stanford’s research site says traditional measures such as lines of code, story points, commit counts, and DORA do not accurately measure engineering productivity. The approaches below can still answer narrower questions, but each has blind spots.

Approach What it measures Work it may miss or distort Safer use
Lines of code or commit counts Activity volume in a code repository. They do not establish usefulness or complexity; they can miss design, mentoring, incident response, and work outside the repository. Operational context, not a standalone individual performance score.
Story points or delivery metrics such as DORA Planning estimates or delivery and operational outcomes, depending on the metric. They do not, by themselves, attribute team results fairly to one person or capture all contributions. Team-level discussion and improvement, interpreted with project context.
Machine-learning code-change review Predicted expert evaluations of code-change dimensions, including complexity and time-related judgments. Work without visible code changes; uncertainty in the prediction and in how a specific score should be interpreted. A prompt for review and investigation, not an automatic employment decision.

These distinctions are especially important when comparing engineers with different responsibilities. A senior engineer may spend substantial time on architecture, review, or mentoring, while another role may involve more direct implementation. A low code-change score can be a useful signal to investigate, but the number alone cannot explain why it is low.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How should a company respond to a low score?

  1. Check what the score represents. Confirm the model’s time window, inputs, threshold, and the kind of code changes it evaluates before treating its output as evidence about a person.
  2. Review the work in context. Ask the engineer and their manager about design, support, review, mentoring, incident response, or other responsibilities that may not be represented in repository data.
  3. Look for corroboration. Compare the signal with project outcomes, peer and manager feedback, role expectations, and relevant work records rather than relying on a single automated score.
  4. Use the finding to diagnose a problem. If evidence points to a performance or workflow issue, investigate causes and agree on clear expectations and support; do not treat a model flag as proof of misconduct or grounds for automatic dismissal.

For enterprise leaders, the practical question is not simply whether an algorithm can flag unusually low visible output. It is whether the organization can validate what the flag means, identify missing context, and act fairly. Without those safeguards, a productivity tool can create false accusations or reward visible activity over useful engineering work.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase07 MAR 2625 minWhat Is a 457 Plan?
  2. The Money DeskBlogTheFinanceBase07 MAR 2621 minTime Value of Money: What It Is and How It Works
  3. The Money DeskBlogTheFinanceBase07 MAR 2627 minAre You Living in One of These Top 10 Most Expensive Cities to Retire?
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.