The 2026 Report · Part 3

Reporting is more likely manipulated than accurate.

The people who assembled the dashboards rated them as more likely to be manipulated than accurate.

By Phil Hatch · Akholi · First Edition, July 2026 · DOI 10.6084/m9.figshare.33005630

The producers’ own verdict

Rated 1–10
How the people who assemble reporting rate itExhibit
246810 Likelihood of manipulationAccuracy of reporting 6.264.81
Source: Akholi Outsourcing 2026. Report producers rated the likelihood that their employer manipulates client-facing dashboards and reports at 6.26 on a 1–10 scale, and rated the accuracy of that same reporting at 4.81 — the people who assemble it rate it more likely manipulated than accurate.
3.1

What the report producers say.

An institution buying agentic AI outsourcing monitors that engagement through reporting the vendor’s own staff assemble. This part examines that channel through the only witnesses available to the study — the people who build the reports. Their account describes a reporting layer with a systemic problem: the people who assembled the dashboards rated them as more likely to be manipulated than accurate, and findings soften at each step up the management chain that interacts with the client. Everything in this part is the vendor workforce describing its own output, with no client verification and no audit.

Asked to rate the likelihood that their employer manipulates the dashboards and reports the client receives, respondents rated it 6.26 out of 10, and 63.61% answered 6 or higher. Respondents rated the accuracy of client-facing reporting at 4.81. Team-average manipulation ratings reach 7 or higher on 56 engagements, and 15 of those remain active today.

Rank softens the reading without materially changing it. Individual Contributors rated the likelihood of report manipulation at 6.39. Engagement Managers, the rank that typically owns the client-facing reports, rated it 5.81. Portfolio Directors rated it 5.59. On average, every rank rated manipulation as more likely than not.

Reported as more common than 5 years ago

Asked whether manipulation is more or less common than 5 years ago, respondents averaged 6.18, against a no-change mark of 5. Most of this workforce is too new to make that comparison, but the finding holds among those who can: respondents with 5 or more years of career history answered 5.64, above the ‘no change’ option.

Morale moves the manipulation reading

The workforce reporting manipulation is also a sign of workforce distress. Respondents who are content and intend to stay rated the manipulation at 4.39; likely leavers rated it 7.41. The correlation between the manipulation rating and job dissatisfaction is r = 0.62 (p < 0.001).

An unhappy workforce can still testify accurately. The mechanism descriptions in 3.2 are specific and consistent.

3.2

Reports shaped by design.

MechanismShare of descriptionsSurvives an accuracy audit
Proxy-metric or favorable-comparator selection27.89%Yes
Exclusion from the sample or denominator18.60%Yes
Permissive completion definitions18.18%Yes
Human-corrected work counted as automated success17.77%Yes
Checkbox and point-in-time self-certification10.33%Partial
Reclassification or blame shift8.88%Partial
Aggregation dilution8.06%Partial
Timing and smoothing5.58%Partial
Hidden queues and side channels5.37%Partial to No
Audit trail suppression5.37%No
Threshold gaming1.86%Partial
Baseline or target gaming1.65%Partial
Outright fabrication1.45%No
Narrative framing1.24%Yes

Table 13 · Reporting mechanisms coded from 484 descriptions

484 respondents across 127 engagements described how client-facing dashboards and reports are manipulated. Their descriptions were independently coded against 14 mechanisms. 4 dominate, and each produces arithmetically correct figures that would likely survive a traditional audit.

27.89% described proxy-metric and favorable-comparator selection — the report carries a metric chosen to flatter. A volume count appears where a correctness rate would be embarrassing, and accuracy is often measured against the agent’s own prior output rather than the client’s actuals.

18.60% described exclusion from the sample or denominator. Failed and awkward cases route to side queues, review buckets, and categories that never enter the published rate; in many cases, poor-performing metrics are excluded from all reporting.

18.18% described permissive completion definitions — the definition of a finished unit of work becomes more lenient as the workload goal is approached.

17.77% described human-corrected work as automated success. Staff correct agent output by hand, and the dashboard counts the corrected result as the agent’s. The human absorption documented in Part 2 is reported to the client as the automation succeeding.

A conventional accuracy audit recomputes the reported numbers and confirms them, and all 4 dominant mechanisms pass it. The defect sits in the choice of metric, the definition of done, and what enters the sample — the arithmetic itself is clean. Catching it requires an audit focused on the exact mechanisms used to generate client-facing reporting.

Mechanism of shaping by the outsourcing model

Outsourcing modelMost describedSecondThird
Customer experience management (n = 133)Permissive completion definitions, 24.06%Human-corrected as automated success, 24.06%Proxy-metric or favorable-comparator selection, 24.06%
Information technology outsourcing (n = 143)Proxy-metric or favorable-comparator selection, 29.37%Exclusion from the sample, 18.18%Permissive completion definitions, 13.99%
Business process outsourcing (n = 122)Proxy-metric or favorable-comparator selection, 26.23%Human-corrected as automated success, 19.67%Exclusion from the sample, 19.67%
Knowledge process outsourcing (n = 86)Proxy-metric or favorable-comparator selection, 33.72%Human-corrected as automated success, 19.77%Checkbox and point-in-time self-certification, 15.12%

Table 14 · The 3 most-described mechanisms by outsourcing model

Shares are of each model’s describers. Third place tied in 2 models: reclassification also reached 13.99% in information technology outsourcing, and permissive completion reached 15.12% in knowledge process outsourcing.

“We tuned the agent’s confidence threshold right before each client review so more breaks got auto-closed as ‘resolved’ that week. We reset it after the meeting.”Individual Contributor, BPO · Canceled engagement
3.3

Reporting softens between the delivery team and the client.

ReadingIndividual ContributorEngagement ManagerPortfolio Director
Rated their own knowledge of the client’s success measures as adequate, among those able to answer35.63%38.67%50.00%
Gave no numeric read on client satisfaction36.58%0%0%
Rating mildness relative to Individual Contributors, all 100 conditionsreference+0.43+1.26
Manipulation likelihood rating, mean6.395.815.59

Table 15 · Line of sight and reading by rank

Across all 100 rated conditions, Engagement Managers rated their engagements 0.43 points better than their own Individual Contributors did, and Portfolio Directors rated them 1.26 points better. The ranks that brief client leadership sit furthest from the daily delivery and carry the most favorable view — reporting travels upward through layers that each see less and rate better than the layer below.

The line-of-sight gap beneath the softening

Much of the delivery organization is unaware of client-defined KPIs or customer satisfaction. 36.58% of Individual Contributors provided no numeric read on whether the client is satisfied; all Engagement Managers and Portfolio Directors provided data on both customer KPIs and customer satisfaction.

Where clarity in customer KPIs exists at the delivery layer, reporting trends toward being accurate — self-assessments closely track the client’s recorded outcomes as reported by respondents. The remainder lack the sight to report against, whatever their intent. A reporting layer that cannot see the client’s definition of success drifts toward the definitions that flatter the vendor.

Continue the report

All parts →
Overview

Report home

The hub: findings, agenda, definitions, and report information.

Read →
Summary

Executive Summary

The sample, the four findings, and the six executive actions.

Read →
Part 1

The Market

Adoption is broad and young, and the losses are already realized.

Read →
Part 2

Why Engagements Fail

Five problem areas; the institutions on both sides come before the technology.

Read →
Part 4

The Delivery Workforce

The oversight layer is junior, dissatisfied, and leaving.

Read →
Part 5

Conclusions & the Executive Agenda

Six actions, all within the institution’s own authority.

Read →
Method

Study Information & AI Disclosure

Design, independence, statistics, limits, and the AI-assistance disclosure in full.

Read →