Study information and AI disclosure.
This page sets out how the study was designed and sampled, the statistical conventions applied throughout, the independence built into the collection, the limits the record carries, and the full disclosure of where artificial intelligence assisted in producing this report.
By Phil Hatch · Akholi · First Edition, July 2026 · DOI 10.6084/m9.figshare.33005630
Study design and sample.
The study recorded structured ratings and written testimony from 850 vendor delivery staff who work on the 150 agentic AI outsourcing engagements at Tier 1 financial institutions. Collection ran throughout the first half of 2026. Respondents hold the 4 vendor delivery roles named throughout the report: Individual Contributors, who execute the work; Engagement Managers, who run a single client engagement; Portfolio Directors, who oversee several Engagement Managers; and Customer Management Executives, who own a client relationship. All 10 primary institution types and all 4 primary outsourcing models are represented.
Recruitment and screening followed a fixed order. A specialist provider built the candidate pool from professional networks, published profiles, and vetted panels, then added referrals from early respondents; participation was compensated, and recruitment reached more than 900 completed responses. Screening removed anyone whose answers showed no working knowledge of banking, the outsourcing industry, or agentic AI, and the screen ran before any analysis, so it could not be tuned toward a preferred result. The retained sample was 850. A validation round followed, during which 122 respondents completed short telephone interviews to confirm their roles and responses, and no material discrepancy appeared.
Two design choices bound what the study can claim. Sampling was purposive rather than proportional: smaller institution types, among them investment banking and wealth management, were oversampled so that each type had enough observations to be described, and engagement counts, therefore, say nothing about market size or share. The account is also vendor-side self-report — systematically instrumented and internally validated, but not corroborated by client records.
The data also hold unresolved internal contradictions, reported as they stand rather than smoothed by assumption. The workforce rates its own confidence against the grain of its experience, and the delivery layer reports both the highest likelihood of manipulation and the least sight of the client measures it is judged against. A market this young, drawn from a small qualified population, produces such patterns, and the significant findings survive them.
Independence by design.
Independence was built into the collection rather than added after it. A 3rd-party research firm administered the survey using its own instrument and retained the raw responses; the author never saw any response in which identifying details might have surfaced. That firm reviewed every response and certified that nothing identifying a bank, a vendor, a bank customer, or a proprietary method had been collected, and the questions were written so that no respondent could name an employer, a client, or a client customer. Both firms behind every engagement stay anonymous by construction.
The anonymity was deliberate. Delivery staff describe their own employer more fully when neither firm can be identified, and the candor of the testimony in Parts 3 and 5 reflects it.
Statistical conventions and precision.
The instrument carried 200 questions, grouped as respondent demographics and sentiment; engagement facts; client and vendor profiles; 100 universal engagement-condition ratings, set in 5 areas of 20; a forced ranking of the 5 issues most limiting each engagement; and open written testimony. Rating scales range from 1 to 10, with 1 being the worst, and a rating of 3 or below is distressed. The answers “I do not know” and “I prefer not to answer” were valid everywhere, and item nonresponse is itself reported as a finding where it occurs. The text quotes ratings to 1 decimal place and shares to whole numbers, and the exhibits keep full precision.
The study carries 2 units of inference, and the report holds them apart. Engagement-level facts — among them age, stage, cancellation, satisfaction, term, and headcount — take one value per engagement and rest on n = 150. Respondent-level readings, among them ratings, sentiment, and testimony, rest on the 850 individual responses; these cluster within engagements, since colleagues on one engagement answer more alike than strangers, with intraclass correlations (ICC) between 0.48 and 0.84 across the rated conditions. Clustering cuts the information in the 850 responses to an effective sample of roughly 170 to 280, and every respondent-level interval below carries that design effect rather than the raw count.
Confidence intervals
The 95% confidence intervals are these. Realized cancellation of 37.33% runs from 30.0% to 45.3%, on n = 150. Mean engagement satisfaction of 43.85% carries about plus or minus 4.5 points, and the active-engagement mean of 64.1% about plus or minus 2.3 points. Workforce job satisfaction of 4.19 carries about plus or minus 0.4 of a point. Departure intent of 53.88% runs from about 47% to 61%. Manipulation rated 6 or higher, at 63.61% among those able to answer, runs from about 58% to 70%. As a working rule, whole-study engagement facts are reliable to within about 8 points, respondent-level shares to within about 7 points, and subgroup readings widen as the cell shrinks.
What counts as a finding
A difference is called a finding only where it survives testing at conventional thresholds, with correction for the number of comparisons. Three cross-cutting results pass. Overall engagement condition is the strongest correlate of cancellation, at r = -0.72; ranked into thirds by condition, the weakest, middle, and strongest thirds cancel at 88%, 24%, and 0%. Engagements on 24-month terms cancel at 56.25%, against 28.00% for all other terms pooled, at a corrected p = 0.010. The age gap between corporate-functional and line-of-business engagements holds at p = 0.003.
The differences the report declines to call findings fail those tests. Cancellation and satisfaction do not differ significantly across institution types, at p = 0.59 or worse, nor across outsourcing models, pricing models, or contract-value bands. Where a cell is small, the report says so, and the 2 Customer Management Executives are reported with their counts and support, not percentages.
Coding the open testimony
The mechanism is shared in Part 3, and the post-mortem themes in Part 2 come from coding the open testimony. 484 manipulation descriptions and 759 advice responses were each coded against a fixed set of categories, with multi-coding allowed. Coded shares move by about 3 points under reasonable changes to the coding rules, and the claims rest on the order and clustering of the stable categories rather than on any single share.
Limitations.
The report carries 7 limitations, stated together. The account is vendor-side self-report, uncorroborated by client records. The portfolio is young, at a median age of 5 months, and was observed once rather than tracked over time. Purposive sampling bars any inference about market composition. Segment cells are too small to separate outcomes by institution type or model. Respondent clustering reduces the effective sample for workforce readings. Forward risk figures are the delivery teams’ own assessments, uncalibrated to realized outcomes. Text-coded shares carry coder sensitivity of about 3 points.
None of these is unusual for a first systematic study of a market this young, and none is hidden in the body. Each finding cites its basis where it is stated.
The report is published for research purposes and does not constitute legal, regulatory, or investment advice.
Artificial intelligence disclosure.
The author prepared this report with the assistance of an artificial intelligence system, Claude from Anthropic. The system worked as a directed tool on defined tasks, and the author reviewed its output. Its role covered 3 areas. It assisted in the statistical analysis of the collected data. It helped structure the report and refine its prose. It checked the dataset and the manuscript to confirm that no bank, no bank client, and no client customer could be identified.
Grammarly was used to address spelling and punctuation throughout the report. Quotations from respondents appear as they were submitted.
The observations, findings, and conclusions in this report are the author’s own. The artificial intelligence does not generate them. Stylistic patterns sometimes read as markers of machine writing. Where they appear here, they are artifacts of Claude’s work in refining the text. The analysis and the judgments remain the author’s.
Continue the report
All parts →Why Engagements Fail
Five problem areas; the institutions on both sides come before the technology.
The Integrity of Performance Information
Reporting the producers rate more likely manipulated than accurate.
Conclusions & the Executive Agenda
Six actions, all within the institution’s own authority.

