Failure concentrates in five problem areas.
Failure concentrates in five problem areas that shift in weight as an engagement matures, and on both sides of the table — the bank’s preparation and the vendor’s delivery capability — the institutions outrank the technology as the source of it.
By Phil Hatch · Akholi · First Edition, July 2026 · DOI 10.6084/m9.figshare.33005630
Where failure concentrates
Ranked by distressed shareBank-side preparation.
Failure across the 150 engagements concentrates in 5 problem areas whose weight shifts as an engagement matures: client and vendor problems dominate in the early stages, and joint co-management leads once engagements reach production. No participant in the study reported an established best practice or proven playbook for client preparation, vendor delivery, or joint co-management. The technology itself is immature and ranks behind institutional failures across every phase, and the 3rd-party ecosystem of tools, data, and services ranks last. The 5 areas move together statistically — an engagement weak in one tends to be weak in the others — so outcomes are attributed to the overall condition of the engagement rather than to any single area.
| Client-readiness group | Group mean rating (1–10) | Rated distressed |
|---|---|---|
| Governance and responsiveness | 3.51 | 51.83% |
| Process readiness | 3.51 | 52.15% |
| People and skills | 3.52 | 51.36% |
| Legacy estate | 3.52 | 52.06% |
| Data readiness | 3.55 | 50.11% |
Table 7 · Client-readiness groups
Client institution conditions were rated the weakest of all the measures in this study. Across all 100 rated conditions, the 5 weakest question groups are those describing the client organization, and more than half of all client-side ratings are categorized as distressed. 5 gaps carry the area, in rated order from weakest to strongest:
Governance and responsiveness rated 3.51. It measures governance, oversight, and the speed of the client’s response, and led to damaging issue selections on 87 of the 150 engagements, with 51.83% rated distressed.
Process readiness rated 3.51. It measures whether the processes handed to agents are documented and mapped at the depth automation requires; documentation gaps led to damaging issue selections on 94 engagements, with 52.15% rated distressed.
People and skills rated 3.52, covering the availability of client staff who understand the processes, the data, and agentic AI. Respondents selected it in 58 engagements, with 51.36% rated distressed.
Legacy estate rated 3.52. It measures the accessibility of core systems built decades before the advent of agentic integration. Respondents selected it in 96 engagements, with 52.06% rated distressed.
Data readiness rated 3.55, covering the quality, normalization, and governance of the data agents run on. Respondents selected it in 93 engagements, with 50.11% rated distressed.
These ratings cluster within a tenth of a point. Client preparation is a uniform condition across everything the engagement depends on, and no single remediation addresses it.
| Lowest-rated client institution conditions | Mean (1–10) | Rated distressed |
|---|---|---|
| Ownership and accountability for agentic AI | 3.47 | 52.00% |
| Stability of the business case, KPIs, and ROI definition | 3.48 | 52.90% |
| Agentic AI knowledge and skills of client staff | 3.48 | 52.71% |
| Willingness to redesign processes for agentic AI | 3.48 | 53.36% |
| Timely access to functional, legacy-system, and process experts | 3.49 | 53.85% |
| Capacity to govern and oversee the risks of AI agents | 3.49 | 52.32% |
| Process mapping and maturity at the depth agentic AI requires | 3.49 | 54.09% |
| Data quality, normalization, and consistency across source systems | 3.50 | 51.05% |
| Data accessibility | 3.51 | 50.38% |
| Legacy infrastructure accessibility, stability, and compatibility | 3.51 | 52.18% |
Table 8 · Lowest-rated client institution conditions
Qualified AI ownership inside the client institution
Clarity over who owns AI within the client institution was rated 3.47, the lowest rating in the study; the stability of the client’s business case and ROI definition was rated 3.48, the second-lowest. Both measure whether the client has decided who is responsible and to what end — an institution that has not settled ownership cannot direct or co-manage an engagement. During working engagements, a clear client-side owner was in place.
Agentic AI knowledge among client staff was rated 3.48, and the capacity to govern AI agent risks was rated 3.49. For canceled engagements, the ownership condition fell to 2.32, from 3.64 among engagements that reached production.
Process maturity, mapping, and optimization for agentic AI
Process mapping and documentation rated 3.49, with 54.09% of ratings distressed. Standardization sat 2 hundredths higher at 3.51, and consistency of process execution across business units followed at 3.56; willingness to redesign processes for agentic AI was the weakest of the 4, at 3.48.
Canceled engagements rated process mapping at 2.38. Engagements that reached production rated it 3.78, a spread of 1.40 points.
Respondents in 47 engagements noted that processes were never re-engineered for agentic AI — an agent executes the process it was given, exactly as mapped, at machine speed. Where the map is missing, the vendor’s team reconstructs the process from observation before any agent can run, and respondents frequently noted that the bank did not validate vendor-mapped or optimized processes; changes were often made without the bank’s awareness.
Client data cleanliness, consistency, and accessibility
Data quality and AI-readiness led to damaging issue selections in 93 of the 150 engagements. Data governance and model-risk readiness followed at 87, data-sharing strain and residency limits appeared in 45, and confidentiality challenges appeared in 38.
Normalization and format consistency across source systems rated 3.50, with 51.05% of ratings distressed. Timely accessibility was rated 3.51, and the quality of the data itself was rated 3.59.
Engagements that reached production rated their data conditions between 3.62 and 3.84, while canceled engagements rated between 2.33 and 2.45. As with process mapping, respondents noted that they are resolving data issues manually, often without the bank’s input or awareness.
Respondents are absorbing the preparation gap
92.00% of Individual Contributors reported manually resolving issues with client-provided materials. Respondents are absorbing the client’s preparation gap through 3 methods of unrecorded work: manual repair of data, legacy access, and other inputs before any agent operates; mapping or enhancing processes to the level agentic AI requires; and defining agentic AI output-quality standards and resolving output errors themselves.
A junior team without material agentic AI or financial-industry experience is making decisions that directly affect output, and most of the vendor’s delivery team lacks a basic understanding of CORAL as it applies to the financial industry. Respondents noted that the client is frequently neither involved nor aware; those decisions are often undocumented and passed down within the vendor team as informal knowledge.
Part of the workforce described this backstop as work it values, naming 4 points of professional pride: catching agent errors before the client sees the output; repairing input data before an agent runs; adjusting the client’s process to fit agentic delivery, without the client’s awareness; and resolving agentic AI output before submission to the bank.
Every prior generation of outsourcing tolerated a lack of client preparation, and human delivery teams absorbed it invisibly. Agentic delivery was sold on removing that human layer. Within these 150 engagements, the human layer takes on a new role — it absorbs the client’s lack of preparation at a speed and volume it was never sized for.
“Tell us who owns this stuff before we go live.”Individual Contributor, BPO · Canceled engagement
Vendor challenges.
| Lowest-rated vendor conditions | Mean (1–10) | Rated distressed |
|---|---|---|
| Depth of knowledge of agentic AI | 3.58 | 51.55% |
| Effectiveness of the employer’s top leadership team | 3.60 | 48.11% |
| Overall qualification to deliver elite agentic AI outsourcing | 3.61 | 47.95% |
| Ability to manage agentic engagements for Tier 1 banks | 3.63 | 48.91% |
| Experience running agentic AI in live production beyond pilots | 3.64 | 48.24% |
| Depth of knowledge of the financial industry | 3.64 | 48.88% |
| Transition from FTE-based to agentic commercial models | 3.65 | 48.57% |
| Playbook from concept to governed production | 3.65 | 48.57% |
| Quality and investment in agentic AI internal training | 3.66 | 48.56% |
| Internal management redesigned for agentic delivery | 3.66 | 48.56% |
Table 9 · Lowest-rated vendor conditions
Issues with the vendor’s own capabilities led to 1,262 damaging-issue selections among all 850 respondents and affected 148 of the 150 engagements — only client unreadiness drew more, at 1,632 selections. Most of the identified issues trace directly to the vendor’s lack of experience with agentic AI.
Entirely new to the vendor
The agentic AI outsourcing industry is less than 2 years old and staffed primarily by people with less than 1 year of experience; most of the vendor-capability challenges respondents noted are rooted in this context. Outsourcing firms are selling agentic delivery to Tier 1 financial institutions with little to no more agentic AI experience or understanding than the banks themselves.
Vendor agentic AI experience separates outcomes. Engagements with vendors that had delivered agentic AI services for 2 years were canceled at a rate of 19.23%; engagements with vendors below that mark were canceled at a rate of over 40%, a difference significant at p = 0.04. Recorded customer satisfaction moves the same way: vendors with 2 years of agentic experience achieved an average satisfaction score of 50.77%, against below 43% for vendors with 1 or fewer years of experience.
Defining the new commercial model under load
The FTE-based billing model is leaving this industry faster than its replacement is arriving. Agentic delivery removes the person from the unit of work, so billing must move from seats supplied to output delivered — a shift vendors in this study have not completed, and are struggling to complete under load.
Senior respondents — the 183 at Engagement Manager rank and above — rated their employer’s remaking of the commercial model for agentic delivery at 4.04, and rated their employer’s ability to price agentic work profitably, now that billable hours no longer measure the work, at 4.02. Respondents noted that the change has created financial pressure within the vendor; every commercial-model indicator in the study ranked between 37.70% and 40.56% distressed.
Weak pricing models directly signal failure and suggest the vendor may be in the early stages of understanding the fundamental changes AI is having on the industry. Pricing capability correlates with cancellation (r = -0.52) and with recorded customer satisfaction (r = 0.62); p < 0.001 for both.
Leadership and management without agentic experience
Vendor leadership teams do not yet hold the experience their firms are selling. Senior respondents rated the overall effectiveness of their employer’s top leadership team at 4.09, against 3.60 across all 850 respondents; 39.89% of the senior ratings on that condition fall in the distressed band. Leadership effectiveness tracks outcomes at the engagement level, correlating at r = -0.51 with cancellation and r = 0.61 with recorded customer satisfaction (p < 0.001 for both). Engagements where senior-leadership reporting was 3 or below were canceled in 64.29% of cases (36 of 56), with recorded satisfaction of 25.75%; where it sat at 6 or higher, cancellation fell to 8.57% (3 of 35), with satisfaction of 67.60%.
Leadership’s understanding of agentic AI widens the gap further. Engagements where leadership’s agentic AI knowledge scored 3 or below canceled at 89.66% (26 of 29) and recorded customer satisfaction of 13.66%. Engagements with a leadership agentic AI rating of 6 or higher recorded no cancellations across all 41 engagements and a mean satisfaction of 71.68%.
The pattern extends to mid-tier management. All 7 direct-manager measures separate canceled from active engagements more sharply than any of the 20 employer-capability conditions. The 5 manager-competence measures correlate with cancellation between r = -0.62 and -0.67 and with recorded satisfaction from r = 0.74 to 0.78, p < 0.001 throughout; the record does not settle the direction of cause, and failing engagements may sour their teams on the manager.
The managers running these engagements reported their own lack of experience with agentic AI. 53.33% of the study’s 150 Engagement Managers reported 0 agentic AI experience, with a mean of 0.47 years; no Engagement Manager reported more than 1 year. Among the 31 Portfolio Directors, 64.52% reported 0 experience, with a mean of 0.33 years. Manager inexperience carries a measurable price: engagements whose Engagement Manager reported 0 agentic years canceled at 45.00% (36 of 80), with recorded satisfaction of 38.44%; engagements whose Engagement Manager reported a single year canceled at 28.57% (20 of 70), with satisfaction of 50.04% — a difference significant at p = 0.04.
The missing vendor-side playbook
Vendors lack a mature agentic AI playbook to govern internal operations and the full customer delivery lifecycle. Respondents rated the extent to which internal management processes have been redesigned for agentic delivery at 3.66, with 48.56% distressed — management running an agentic business through a pre-agentic internal process, with no firm-level rules to inherit.
Respondents rated their employer’s ability to manage agentic engagements end-to-end at 3.69, and rated the maturity of the playbook for taking agentic use cases from concept to governed production at 3.65, with 48.57% of ratings in the distressed range. On canceled engagements that reading fell to 2.40; engagements in production stood at 3.94. The absence of a mature playbook led to damaging issue selections in 111 of the 150 engagements. Playbook maturity correlates with cancellation at r = -0.59 and with recorded customer satisfaction at r = 0.71 (p < 0.001 for both).
Vendors are not maturing their customer-facing playbooks fast enough. Among the 94 active engagements, playbook maturity is flat against engagement age: 4.47 up to 3 months old, 4.34 at 4 to 6 months, and 4.45 beyond 7 months. Respondents noted that without an operationally mature client-facing playbook, they invent the rules themselves — junior staff making delivery decisions without formal guidance from their employer, without direct customer involvement, and often outside the visibility of either party’s management team.
The vendor’s own platforms predate agents
Over the last decade, outsourcing vendors have expanded hybrid outsourcing — bundling cloud, SaaS, and other platforms with the human delivery layer — and respondents described these platforms as problematic in an agentic AI context. Respondents rated how well their employer’s bundled legacy platforms are optimized for agentic delivery at 3.70, with 48.36% of ratings in the distressed range; on canceled engagements the reading fell to 2.52. Respondents selected vendor-built platforms that are not agent-ready as the single most damaging issue on 26 engagements.
No agentic talent exists for the vendor to hire
Respondents rated their employer’s supply of qualified agentic AI talent against realized demand at 3.69, with 47.63% of ratings in the distressed range. The rating does not describe a hiring backlog — there is no pool of experienced practitioners for any vendor to recruit, and no vendor in the study had solved the agentic AI talent shortage.
The shortage reached the engagements directly. Respondents selected talent shortage as one of the 5 most damaging issues in 69 of the 150 engagements, and selections concentrated in failure: 57.14% of canceled engagements had a talent-shortage selection, against 39.36% of active ones. Canceled engagements rated the talent supply at 2.46, against 3.93 in production.
With no talent to hire, the vendor’s only remaining lever is training, and respondents rated that as problematic too. Employer training programs for practical agentic AI skills rated 3.69, and for financial-industry knowledge 3.72; on canceled engagements the two readings fell to 2.50 and 2.57. Of the 800 respondents who named a primary source of agentic AI knowledge, only 125 named their employer’s training program (15.62% share) — last among the 6 sources measured, behind vendor and partner training (140), online courses and on-the-job learning (136 each), university coursework (132), and self-teaching (131).
“I’m a fresher and know more about AI than anyone in our executive team.”Individual Contributor, KPO · Active engagement
Co-management.
| Lowest-rated co-management conditions | Mean (1–10) | Rated distressed |
|---|---|---|
| Readiness of the joint management model for agent autonomy | 3.60 | 50.37% |
| Quality measures capture the errors that matter | 3.61 | 51.14% |
| Clarity on what agents may do without human approval | 3.61 | 50.25% |
| Speed to jointly stop an agent behaving unexpectedly | 3.62 | 50.50% |
| Agent updates are visible to the other firm before they take effect | 3.62 | 49.07% |
| The governance model fits the work performed by agents | 3.63 | 49.19% |
| Single shared view of the risks agents introduce | 3.64 | 48.30% |
| Human review points placed where errors matter most | 3.64 | 49.31% |
| Genuine scrutiny of agent output at review points | 3.65 | 49.50% |
| Joint issue resolution matches agent speed | 3.65 | 49.44% |
Table 10 · Lowest-rated co-management conditions
Joint daily management of an engagement whose work is executed by autonomous systems has no precedent in outsourcing practice, and the methods and maturity of joint customer-vendor delivery have yet to be developed in an agentic AI context. Traditional outsourcing governance was paced to people — review meetings, issue resolution, and change boards assume work at human speed. Agentic delivery moves faster than any of those processes and arrives in volumes they cannot clear, and neither the client institutions nor the vendors in this study have built replacements; the market is too young to have produced best practices to adopt.
Among production-stage engagements, 3 conditions rated lowest. Cross-firm visibility of agent actions was rated 3.25, the lowest among the production engagement conditions. Joint human review capacity rated 3.26 — a review built for human throughput cannot clear machine throughput. Joint change control rated 3.3: agents change behavior between formal releases as a model updates, a prompt changes, or upstream data shifts, and respondents frequently make material changes to ensure KPIs are met; a change board built for quarterly release cycles governs none of it, and respondents reported changes reaching production without any joint review.
No co-management playbook exists to adopt. Engagements that survive are those in which the client and vendor build one together before the contract is signed, and the highest-rated engagements keep refining it throughout the engagement.
Co-management becomes critical as the engagement matures
Early-stage engagements rated client-readiness conditions as their weakest category. Among production engagements and all canceled engagements, the 15 weakest-rated of the 100 conditions are all co-management conditions — client and vendor problems are worked down as an engagement matures, and the co-management problem becomes the most impactful category. The 10 lowest-rated co-management conditions range from 3.60 to 3.65 across the study, a spread of 0.5; limited to production engagements, the 3 weakest co-management conditions were rated between 3.25 and 3.27.
Decision rights and accountability
The weakest condition among canceled engagements is clarity on what agents may do without human approval, rated 2.16. Decision rights over agent autonomy are the first question a joint agentic operation must answer, and no canceled engagement had settled it. Severity rises as the rating moves down the management chain toward the daily work: Individual Contributors rated co-management at 3.52 against 4.98 among Portfolio Directors on the same engagements, a gap of 1.46 points.
Cross-firm visibility of agent actions
Agent visibility across the firm boundary was rated 3.25 among production engagements. Asked to name the single most damaging issue in their engagement, respondents chose the agent black box across the boundary 54 times. Visibility of the other party’s agent updates was rated 3.29. Canceled engagements rated the same condition at 2.3; genuine scrutiny of agent output rated 2.16 on canceled engagements, and the quality measures both firms used to validate the other party’s agents rated 2.21. Production engagements rated cross-firm visibility at 3.9, 1.6 points above canceled engagements.
Human review at machine output speed
Joint human review capacity was rated 3.26 across production engagements and led to damaging issue selections in 63 engagements. Placement of review points where errors matter most was rated 3.30, and genuine scrutiny of work product at agentic AI output scale rated 3.4. The study indicates that much of the human review process still operates in a pre-AI context; given non-deterministic AI outputs, humans need to review the work, but the review process must be optimized and scaled specifically for agentic AI.
Joint change control at machine speed
Respondents noted that they must constantly effect changes in response to variations in inputs, customer demands, 3rd-party systems, and the non-deterministic output of AI; the volume of changes teams make exceeds the capability of the customer-vendor change control process, against high demands to achieve KPIs.
Joint change control rated 3.3 among production engagements. The speed mismatch beneath it was measured directly: joint change control processes kept pace with the speed and volume of change at a rating of 3.4 in production, against 2.26 for canceled work. Canceled engagements rated joint governance of human- or agent-made changes at 2.19, against 3.32 in production.
“We were so busy fixing pre-prod issues that we never thought about how to do the actual work.”Engagement Manager, KPO · Active engagement
The technology ranks behind the institutions.
| Lowest-rated technology conditions | Mean (1–10) | Rated as distressed |
|---|---|---|
| Production readiness of the tools for bank-grade work | 4.54 | 32.04% |
| Safe handling of sensitive data across outputs and logs | 4.54 | 32.38% |
| End-to-end reliability on multi-step tasks | 4.55 | 32.48% |
| Accuracy of citations and references against real sources | 4.55 | 33.50% |
| Agents took only intended and authorized actions | 4.56 | 32.07% |
| Support for reliable pre-release testing | 4.56 | 32.70% |
| Resistance to hostile instructions in processed content | 4.56 | 32.14% |
| Integrity of persistent memory against contamination | 4.57 | 31.89% |
| Ability to explain and provide evidence outputs at audit depth | 4.57 | 32.75% |
| Stability of behavior when nothing changed | 4.57 | 33.00% |
Table 11 · Lowest-rated technology conditions
Inside these engagements, agentic AI technology does not meet the standards the financial industry demands. A third of all technology ratings sit in the distressed band, and output-quality defects — including fabricated and inaccurate content — appear throughout the written testimony as a routine operating burden. For institutions whose tolerance for error in regulated processes approaches 0, a technology-layer rating of 4.54–4.66 falls well short of the bank-grade standard. The challenges with the core technology are widely documented beyond this study.
Non-repeating AI output and hallucinations
Agentic AI output is non-deterministic: the same agent, given identical input and configuration, returns different outputs across runs. Respondents rated the tools’ consistency under identical inputs at 4.58. No vendor in the study carried a proven method for evaluating non-repeating output; employer QA capability for such output rated 3.68, with 47.48% of ratings distressed. QA and evaluation gaps led to damaging issue selections in 111 of the 150 engagements, matching the breadth of the missing delivery playbook — an engagement without a playbook improvises its delivery, and an engagement without an evaluation method cannot prove what the improvisation produced.
Teams re-run outputs and compare versions by hand, maintaining non-AI systems in parallel with the agents they operate to validate output and certify samples. None of that labor appears in any client-facing measure. The effort is the cost of operating a non-deterministic system under a deterministic quality model, and the data indicate the vendor absorbs it silently.
“AI invented a husband for a widow on her loan application. I looked at the processing logs. I don’t know why it thought this was important, or how it created this new person.”Portfolio Director, BPO · Active engagement
The third-party ecosystem ranks last.
| Lowest-rated 3rd-party conditions | Mean (1–10) | Rated distressed |
|---|---|---|
| Deliberate handoff of the quiet human fixes for 3rd-party quirks | 4.58 | 33.54% |
| Change notice reaching the people running the agents | 4.58 | 31.95% |
| Managing dependence on a small number of providers | 4.59 | 32.04% |
| Rate limits and quotas under agent volume | 4.60 | 32.92% |
| Licenses covering work done by agents rather than named users | 4.61 | 31.91% |
| Knowing where humans quietly fixed 3rd-party issues | 4.61 | 33.08% |
| Agent dependence on gated external rails | 4.61 | 30.86% |
| Fallback readiness for a critical 3rd-party failure | 4.61 | 32.04% |
Table 12 · Lowest-rated 3rd-party conditions
Respondents rated the 3rd-party ecosystem the least severe of the 5 problem areas. Failures inside it remain material and reach every engagement.
Change and service disruption notifications
Change arrives from third parties unannounced — a model version retires, or an API shifts underneath running agents. Upstream data structures change on a cadence the providers alone control, and they owe the engagement nothing. Change notifications and service interruptions are not reaching the delivery team: notice quality rated 4.6, and the paths that would carry notice to the agents’ operators rated the same.
Programmatic access
Respondents rated the availability of effective machine interfaces at 4.63, with 29.80% of ratings distressed. A machine interface replaces a screen built for people. Where APIs exist, their stability was rated 4.66, and the quality of the data arriving through them was rated 4.64. Access to gated external rails, including payment networks and regulatory submission systems, was rated 4.61.
Production engagements rated 3rd-party machine-interface availability at 5.04; canceled engagements rated it 3.14, among the widest spreads in the 3rd-party area. The exposure sits in the ecosystem itself — systems built for human operators now run under agents that cannot use them. Survival and programmatic access move together in this record, whichever drives the other.
“No API, no way in.”Individual Contributor, CXM · Canceled engagement
The five problem areas across the lifecycle.
An engagement opens against a client institution that has not prepared adequately for agentic AI and has not settled ownership. Delivery then falls to a vendor with at most 2 years of experience doing the work, depending on a workforce without any meaningful agentic AI experience. An engagement that survives those 2 conditions long enough to mature meets the third: no one in this study has built an effective joint co-management discipline. Underneath sits an immature technology, and around it an ecosystem that has not yet evolved for agentic use.
Failure in this study set is institutional in origin rather than technical. Those conditions sit within the buying institution’s authority to change. Part 5 defines the resolution actions.
Continue the report
All parts →The Integrity of Performance Information
Reporting the producers rate more likely manipulated than accurate.
Conclusions & the Executive Agenda
Six actions, all within the institution’s own authority.
Study Information & AI Disclosure
Design, independence, statistics, limits, and the AI-assistance disclosure in full.

