Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Upwork’s initial Human+Agent Productivity Index found that expert feedback raised AI agents’ completion rates by up to 70% on a selected set of real projects. That is a relative improvement claim—not necessarily a 70-percentage-point increase or a productivity gain. The benchmark covered 322 low-complexity jobs, not professional work in general, and counted a job as complete only when an output met every evaluator-defined criterion.
What Upwork found
Upwork announced the initial findings on November 13, 2025. The company reported that human feedback increased completion rates by up to 70% compared with agents working alone. The public headline does not establish that this was the average gain across all jobs, nor does it supply a single overall baseline and improved rate. It should therefore be read as Upwork’s maximum comparative claim, not as a universal result. Upwork’s announcement
The benchmark’s initial dataset comprised 322 low-complexity jobs that had previously been posted, paid for and successfully completed by verified Upwork clients and freelancers. Upwork says those jobs represented less than 6% of its gross services volume. Ninety percent of the project budgets were between $10 and $200; project durations ranged from about nine hours to more than 100 days, so a low budget did not necessarily mean a short project. Upwork’s HAPI methodology
How the benchmark worked
Real projects, selected for clear scope
The jobs spanned accounting and consulting, administrative support, data science and analytics, engineering and architecture, sales and marketing, translation, web, mobile and software development, and writing. They had defined scopes, milestones and requirements. Upwork excluded jobs with multiple milestones, price changes or personally identifiable information.
#1 Best Overall
These were real marketplace projects, but they were deliberately bounded and relatively favorable conditions for agents: the work had a clear specification and a successful human-completed outcome to refer to. Upwork says it excluded open-ended and highly complex projects, which it describes as typical of the vast majority of work on its platform. The result does not represent a random sample of every kind of freelance work, including abandoned or unsuccessful projects. Upwork’s methodology
Completion meant passing every criterion
Experienced freelancers created a task-specific rubric of five to 20 pass/fail criteria. Upwork says evaluators had 100% Job Success Scores and Top Rated or Top Rated Plus status; collectively, they had completed more than 96,000 hours of work and earned more than $1 million on the platform. The benchmark measured the share of jobs for which an agent met all criteria, first on its own and then after cycles of human feedback.
That is a demanding all-or-nothing threshold: missing one criterion means the job does not count as complete. But passing a rubric is not the same as proving a client would pay for, publish, deploy or accept the result without further revision. Rubric completion does not by itself establish originality, persuasive force, legal safety or commercial usefulness. The score describes performance against the defined criteria, not every dimension of professional quality.
Models and reported examples
Upwork’s public methodology page gives the broad design and result, but not a full model-by-category results table. VentureBeat reported that the systems tested included Google Gemini 2.5 Pro, OpenAI GPT-5 and Anthropic Claude Sonnet 4. Its account gives these examples; they should be treated as secondary reporting, not as a complete or independently audited table of HAPI results:
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →| Reported example | Working alone | After feedback |
|---|---|---|
| Claude Sonnet 4, data science and analytics | 64% | 93% |
| Gemini 2.5 Pro, sales and marketing | 17% | 31% |
| GPT-5, engineering and architecture | 30% | 50% |
| Claude Sonnet 4, web development | 68% | Not specified in VentureBeat’s report |
| Gemini 2.5 Pro, selected technical tasks | Up to 74% | Not specified in VentureBeat’s report |
Source for these model-specific figures: VentureBeat’s account of the study. They illustrate variation by task and model; they do not establish a single rate for all agents or projects.
Where the results point—and where they do not
Structured technical work was a relative strength
Upwork says agents performed best in structured technical categories such as coding and data science. Human expertise still improved results across technical work, including web, mobile and software development. A strong score in a bounded technical task does not mean an agent can independently manage a complete software project, where requirements, integration, security and changing priorities may matter.
Context and judgment are harder to reduce to a checklist
Writing, translation, sales and marketing, and some engineering and architecture tasks can depend on taste, cultural nuance, business context and unstated client intent. Several answers may be defensible, and “good enough” may depend on the audience or objective. VentureBeat reported especially large feedback gains in some qualitative categories, but the official public summary does not provide the full category-by-model breakdown needed to generalize those examples.
“Fail independently” does not mean “do no useful work”
On this benchmark, an agent working alone could produce useful or partly correct work and still fail the completion measure by missing one or more rubric criteria. The result is about reliability against a strict acceptance threshold without feedback; it does not show that every agent failed every task. Nor does a missed criterion reveal on its own whether the cause was weak capability, ambiguity in the job, an incomplete rubric or a missing piece of context.
Best Value
Why real marketplace tasks matter—and why the test remains limited
Static benchmarks can test a controlled capability, but they do not necessarily reproduce the gap between a clean prompt and a client request with implicit expectations, revisions and commercial consequences. HAPI’s use of completed marketplace jobs and task-specific criteria is a useful attempt to bring evaluation closer to professional work. Upwork’s associated UpBench paper describes a benchmark approach grounded in verified jobs and expert rubrics; it supports the methodological direction, but is not independent confirmation of every HAPI result. The UpBench paper
Several limits matter when applying the findings:
- Restricted task mix: the initial sample emphasized simple, fixed-price, clearly scoped work and excluded much more complex activity.
- Rubric coverage: a checklist can capture stated requirements while missing originality, professional judgment or whether a deliverable works in its real business setting.
- Human contribution: feedback may clarify requirements, add domain knowledge, change task decomposition, identify errors or contribute substantive work. The result evaluates a model-plus-human workflow, not model capability in isolation.
- Feedback details: VentureBeat reported that evaluators spent about 20 minutes per review cycle, but that figure is not prominent in Upwork’s public summary. The public materials do not establish enough detail to treat the feedback process as a universal labor estimate.
- Commercial incentives: Upwork designed the benchmark, used its marketplace data and selected the evaluators. The company has a strategic interest in showing the value of human expertise alongside AI. That does not invalidate the results, but makes them useful early evidence rather than a neutral final verdict.
- Changing systems: model versions, tools, browsing, code execution and access to external systems can change performance. These results do not establish how newer or differently configured agents perform.
What the findings mean for businesses
The strongest practical implication is to evaluate an agent as part of a workflow. The relevant question is not only whether it can produce a plausible first answer, but how much skilled human time is needed to turn that answer into an accepted result—and how costly an error would be.
A supervised workflow
- Define the objective and acceptance criteria. A human should state what success looks like, what constraints apply and which decisions require judgment.
- Ask the agent for a first pass. Keep the task bounded and provide the necessary context and permitted data.
- Review the output. Check factual accuracy, assumptions, constraints, completeness and fit for the intended audience or system.
- Give specific feedback. Identify the missed requirement or error and explain what needs to change, rather than asking vaguely for a better answer.
- Have the agent revise, then obtain human approval. A qualified person should decide whether the result is fit to use and remain accountable for consequential decisions.
- Automate further only after measuring repeated performance. Confirm that the workflow is stable across representative tasks, not just one successful example.
Match oversight to the task
- Consider lighter supervision for repetitive, well-defined work with objective checks, structured data, reversible outcomes and low-cost errors.
- Keep a domain expert closely involved when instructions are ambiguous, client taste or cultural context matters, mistakes are hard to detect, or the work affects revenue, reputation, compliance or safety.
- Do not delegate sensitive material casually. Work involving confidential or regulated information needs appropriate privacy and security controls, regardless of whether a human or agent produces the deliverable.
Track accepted work, not just generated answers
For each task type, record first-pass completion, human intervention count, review time, revision count, error severity, cost per accepted output, time to final acceptance, client rejection and escalation rates. Those measures show whether an agent actually reduces labor or merely moves it into review and repair. Upwork’s reported completion-rate increase alone does not establish a 70% productivity, revenue or profit gain.
What the study does not prove
HAPI does not prove that people will always need to review every task, that agents cannot become reliable independently, or that human review is always cheaper. It also does not show that every human contribution improves the result or that freelancers necessarily benefit economically from AI-enabled work. Its narrower evidence is that, on Upwork’s selected set of real, relatively simple projects, human feedback improved the share of jobs meeting every defined criterion.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




