Not conclusively. A Stanford-linked research effort estimated that 9.5% of software engineers in its dataset produced less than one tenth of the median engineer’s measured output. That is a striking early finding, not a settled estimate of how many engineers do little work or a basis for calling them a drain on their employers. The model looks at code changes, and important engineering contributions may not appear in them.
What does the “ghost engineer” claim mean?
In 2024, researcher Yegor Denisov-Blanch described engineers in the lowest-output category as “0.1x-ers,” saying they did “virtually nothing.” ITPro reported the estimate as 9.5% of software engineers and said the dataset covered more than 50,000 engineers at hundreds of companies. The claim concerns measured output relative to the median in that dataset; it does not establish that those engineers performed no work.
The distinction matters. A relative score from one dataset is not automatically a representative estimate for the software workforce as a whole. The figure should be read as a result reported by one research effort, rather than a universal industry statistic or a verified count of employees who contribute nothing.
How does the productivity model work?
The Stanford Software Engineering Productivity Research site describes an approach that uses machine learning to approximate expert evaluations of software commits. Denisov-Blanch has said the system analyzes source-code changes in private Git repositories and simulates a panel of 10 expert reviewers. He has also clarified that it is intended to assess the substance and complexity of changes, not simply count commits or lines of code.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11The related 2024 arXiv preprint, Predicting Expert Evaluations in Software Code Reviews, reports correlations of r = 0.82 for coding time and r = 0.86 for implementation time between model predictions and expert evaluations. Those figures describe agreement on particular review dimensions; they do not show that the model accurately identifies every low-performing employee or establishes the 9.5% estimate as a population-level fact. The abstract says the work addresses review dimensions “typically avoided due to their complexity or subjectivity.”
What did the study report about remote work?
Denisov-Blanch’s 2024 public thread reported different shares of engineers in the “ghost” category by work location:
Rank #2
| Work arrangement | Reported share | How to interpret it |
|---|---|---|
| Fully remote | 14% | Author-reported subgroup estimate from the same research effort; not an independently established prevalence rate. |
| Hybrid | 9% | Author-reported subgroup estimate from the same research effort; not an independently established prevalence rate. |
| Office-based | 6% | Author-reported subgroup estimate from the same research effort; not an independently established prevalence rate. |
The figures suggest a difference within the reported dataset, but they do not show that remote work causes lower productivity. The public account does not establish that the groups were comparable in role, seniority, company, project, or other conditions. The percentages therefore should not be used to rank work arrangements or predict an individual engineer’s performance.
Why is the 9.5% figure contested?
Jellyfish noted that the viral estimate had not been peer reviewed and that neither the paper nor the researcher’s post clearly explained how the 9.5% figure was derived. Pluralsight raised questions about how a larger dataset—1.73 million commits and 50,935 engineers—related to model training. It also argued that commit-centered analysis can overlook mentoring, debugging, architecture, code review, and other work.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #3
These critiques do not prove the estimate wrong. They identify limits on what can be concluded from the public account: the method behind the specific threshold and estimate is not described clearly enough there to treat the result as independently validated workforce prevalence. The dataset figures reported by ITPro and discussed in the critique also should not be assumed to describe the same analytical stage or sample without a clear explanation of that relationship.
Denisov-Blanch has said participating organizations checked many flagged engineers and that ancillary activities did not explain most cases. That is the author’s account, not independent confirmation of the estimate. He has also cautioned that “Decisions shouldn’t be made purely based on what our model spits out.”
Rank #4
What can different productivity measures actually tell an employer?
No single metric captures every part of engineering work. Stanford’s research site says traditional measures such as lines of code, story points, commit counts, and DORA do not accurately measure engineering productivity. The approaches below can still answer narrower questions, but each has blind spots.
| Approach | What it measures | Work it may miss or distort | Safer use |
|---|---|---|---|
| Lines of code or commit counts | Activity volume in a code repository. | They do not establish usefulness or complexity; they can miss design, mentoring, incident response, and work outside the repository. | Operational context, not a standalone individual performance score. |
| Story points or delivery metrics such as DORA | Planning estimates or delivery and operational outcomes, depending on the metric. | They do not, by themselves, attribute team results fairly to one person or capture all contributions. | Team-level discussion and improvement, interpreted with project context. |
| Machine-learning code-change review | Predicted expert evaluations of code-change dimensions, including complexity and time-related judgments. | Work without visible code changes; uncertainty in the prediction and in how a specific score should be interpreted. | A prompt for review and investigation, not an automatic employment decision. |
These distinctions are especially important when comparing engineers with different responsibilities. A senior engineer may spend substantial time on architecture, review, or mentoring, while another role may involve more direct implementation. A low code-change score can be a useful signal to investigate, but the number alone cannot explain why it is low.
Recommended Free Tools
Best Value
How should a company respond to a low score?
- Check what the score represents. Confirm the model’s time window, inputs, threshold, and the kind of code changes it evaluates before treating its output as evidence about a person.
- Review the work in context. Ask the engineer and their manager about design, support, review, mentoring, incident response, or other responsibilities that may not be represented in repository data.
- Look for corroboration. Compare the signal with project outcomes, peer and manager feedback, role expectations, and relevant work records rather than relying on a single automated score.
- Use the finding to diagnose a problem. If evidence points to a performance or workflow issue, investigate causes and agree on clear expectations and support; do not treat a model flag as proof of misconduct or grounds for automatic dismissal.
For enterprise leaders, the practical question is not simply whether an algorithm can flag unusually low visible output. It is whether the organization can validate what the flag means, identify missing context, and act fairly. Without those safeguards, a productivity tool can create false accusations or reward visible activity over useful engineering work.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




