The 2026 Report · Part 2

Failure concentrates in five problem areas.

Failure concentrates in five problem areas that shift in weight as an engagement matures, and on both sides of the table — the bank’s preparation and the vendor’s delivery capability — the institutions outrank the technology as the source of it.

By Phil Hatch · Akholi · First Edition, July 2026 · DOI 10.6084/m9.figshare.33005630

Where failure concentrates

Ranked by distressed share
The five problem areas, ranked by share of distressed ratingsExhibit
20%40%60% Bank-side (client) preparationBank–vendor co-managementVendor preparation & mgmtAI tools (technology)3rd-party ecosystem 3.523.653.664.584.62 51.48%49.26%48.22%32.14%31.74%
Source: Akholi Outsourcing 2026, Table 6 (n = 850). Bars show the share of ratings in the distressed band (rated 3 or below). The figure inside each bar is the mean condition score on a 1–10 scale, where 1 is worst and 10 is best; the institution- and relationship-side areas (amber) score materially worse than technology and the third-party ecosystem (grey).
2.1

Bank-side preparation.

Failure across the 150 engagements concentrates in 5 problem areas whose weight shifts as an engagement matures: client and vendor problems dominate in the early stages, and joint co-management leads once engagements reach production. No participant in the study reported an established best practice or proven playbook for client preparation, vendor delivery, or joint co-management. The technology itself is immature and ranks behind institutional failures across every phase, and the 3rd-party ecosystem of tools, data, and services ranks last. The 5 areas move together statistically — an engagement weak in one tends to be weak in the others — so outcomes are attributed to the overall condition of the engagement rather than to any single area.

Client-readiness groupGroup mean rating (1–10)Rated distressed
Governance and responsiveness3.5151.83%
Process readiness3.5152.15%
People and skills3.5251.36%
Legacy estate3.5252.06%
Data readiness3.5550.11%

Table 7 · Client-readiness groups

Client institution conditions were rated the weakest of all the measures in this study. Across all 100 rated conditions, the 5 weakest question groups are those describing the client organization, and more than half of all client-side ratings are categorized as distressed. 5 gaps carry the area, in rated order from weakest to strongest:

Governance and responsiveness rated 3.51. It measures governance, oversight, and the speed of the client’s response, and led to damaging issue selections on 87 of the 150 engagements, with 51.83% rated distressed.

Process readiness rated 3.51. It measures whether the processes handed to agents are documented and mapped at the depth automation requires; documentation gaps led to damaging issue selections on 94 engagements, with 52.15% rated distressed.

People and skills rated 3.52, covering the availability of client staff who understand the processes, the data, and agentic AI. Respondents selected it in 58 engagements, with 51.36% rated distressed.

Legacy estate rated 3.52. It measures the accessibility of core systems built decades before the advent of agentic integration. Respondents selected it in 96 engagements, with 52.06% rated distressed.

Data readiness rated 3.55, covering the quality, normalization, and governance of the data agents run on. Respondents selected it in 93 engagements, with 50.11% rated distressed.

These ratings cluster within a tenth of a point. Client preparation is a uniform condition across everything the engagement depends on, and no single remediation addresses it.

Lowest-rated client institution conditionsMean (1–10)Rated distressed
Ownership and accountability for agentic AI3.4752.00%
Stability of the business case, KPIs, and ROI definition3.4852.90%
Agentic AI knowledge and skills of client staff3.4852.71%
Willingness to redesign processes for agentic AI3.4853.36%
Timely access to functional, legacy-system, and process experts3.4953.85%
Capacity to govern and oversee the risks of AI agents3.4952.32%
Process mapping and maturity at the depth agentic AI requires3.4954.09%
Data quality, normalization, and consistency across source systems3.5051.05%
Data accessibility3.5150.38%
Legacy infrastructure accessibility, stability, and compatibility3.5152.18%

Table 8 · Lowest-rated client institution conditions

Qualified AI ownership inside the client institution

Clarity over who owns AI within the client institution was rated 3.47, the lowest rating in the study; the stability of the client’s business case and ROI definition was rated 3.48, the second-lowest. Both measure whether the client has decided who is responsible and to what end — an institution that has not settled ownership cannot direct or co-manage an engagement. During working engagements, a clear client-side owner was in place.

Agentic AI knowledge among client staff was rated 3.48, and the capacity to govern AI agent risks was rated 3.49. For canceled engagements, the ownership condition fell to 2.32, from 3.64 among engagements that reached production.

Process maturity, mapping, and optimization for agentic AI

Process mapping and documentation rated 3.49, with 54.09% of ratings distressed. Standardization sat 2 hundredths higher at 3.51, and consistency of process execution across business units followed at 3.56; willingness to redesign processes for agentic AI was the weakest of the 4, at 3.48.

Canceled engagements rated process mapping at 2.38. Engagements that reached production rated it 3.78, a spread of 1.40 points.

Respondents in 47 engagements noted that processes were never re-engineered for agentic AI — an agent executes the process it was given, exactly as mapped, at machine speed. Where the map is missing, the vendor’s team reconstructs the process from observation before any agent can run, and respondents frequently noted that the bank did not validate vendor-mapped or optimized processes; changes were often made without the bank’s awareness.

Client data cleanliness, consistency, and accessibility

Data quality and AI-readiness led to damaging issue selections in 93 of the 150 engagements. Data governance and model-risk readiness followed at 87, data-sharing strain and residency limits appeared in 45, and confidentiality challenges appeared in 38.

Normalization and format consistency across source systems rated 3.50, with 51.05% of ratings distressed. Timely accessibility was rated 3.51, and the quality of the data itself was rated 3.59.

Engagements that reached production rated their data conditions between 3.62 and 3.84, while canceled engagements rated between 2.33 and 2.45. As with process mapping, respondents noted that they are resolving data issues manually, often without the bank’s input or awareness.

Respondents are absorbing the preparation gap

92.00% of Individual Contributors reported manually resolving issues with client-provided materials. Respondents are absorbing the client’s preparation gap through 3 methods of unrecorded work: manual repair of data, legacy access, and other inputs before any agent operates; mapping or enhancing processes to the level agentic AI requires; and defining agentic AI output-quality standards and resolving output errors themselves.

A junior team without material agentic AI or financial-industry experience is making decisions that directly affect output, and most of the vendor’s delivery team lacks a basic understanding of CORAL as it applies to the financial industry. Respondents noted that the client is frequently neither involved nor aware; those decisions are often undocumented and passed down within the vendor team as informal knowledge.

Part of the workforce described this backstop as work it values, naming 4 points of professional pride: catching agent errors before the client sees the output; repairing input data before an agent runs; adjusting the client’s process to fit agentic delivery, without the client’s awareness; and resolving agentic AI output before submission to the bank.

Every prior generation of outsourcing tolerated a lack of client preparation, and human delivery teams absorbed it invisibly. Agentic delivery was sold on removing that human layer. Within these 150 engagements, the human layer takes on a new role — it absorbs the client’s lack of preparation at a speed and volume it was never sized for.

“Tell us who owns this stuff before we go live.”Individual Contributor, BPO · Canceled engagement
2.2

Vendor challenges.

Lowest-rated vendor conditionsMean (1–10)Rated distressed
Depth of knowledge of agentic AI3.5851.55%
Effectiveness of the employer’s top leadership team3.6048.11%
Overall qualification to deliver elite agentic AI outsourcing3.6147.95%
Ability to manage agentic engagements for Tier 1 banks3.6348.91%
Experience running agentic AI in live production beyond pilots3.6448.24%
Depth of knowledge of the financial industry3.6448.88%
Transition from FTE-based to agentic commercial models3.6548.57%
Playbook from concept to governed production3.6548.57%
Quality and investment in agentic AI internal training3.6648.56%
Internal management redesigned for agentic delivery3.6648.56%

Table 9 · Lowest-rated vendor conditions

Issues with the vendor’s own capabilities led to 1,262 damaging-issue selections among all 850 respondents and affected 148 of the 150 engagements — only client unreadiness drew more, at 1,632 selections. Most of the identified issues trace directly to the vendor’s lack of experience with agentic AI.

Entirely new to the vendor

The agentic AI outsourcing industry is less than 2 years old and staffed primarily by people with less than 1 year of experience; most of the vendor-capability challenges respondents noted are rooted in this context. Outsourcing firms are selling agentic delivery to Tier 1 financial institutions with little to no more agentic AI experience or understanding than the banks themselves.

Vendor agentic AI experience separates outcomes. Engagements with vendors that had delivered agentic AI services for 2 years were canceled at a rate of 19.23%; engagements with vendors below that mark were canceled at a rate of over 40%, a difference significant at p = 0.04. Recorded customer satisfaction moves the same way: vendors with 2 years of agentic experience achieved an average satisfaction score of 50.77%, against below 43% for vendors with 1 or fewer years of experience.

Defining the new commercial model under load

The FTE-based billing model is leaving this industry faster than its replacement is arriving. Agentic delivery removes the person from the unit of work, so billing must move from seats supplied to output delivered — a shift vendors in this study have not completed, and are struggling to complete under load.

Senior respondents — the 183 at Engagement Manager rank and above — rated their employer’s remaking of the commercial model for agentic delivery at 4.04, and rated their employer’s ability to price agentic work profitably, now that billable hours no longer measure the work, at 4.02. Respondents noted that the change has created financial pressure within the vendor; every commercial-model indicator in the study ranked between 37.70% and 40.56% distressed.

Weak pricing models directly signal failure and suggest the vendor may be in the early stages of understanding the fundamental changes AI is having on the industry. Pricing capability correlates with cancellation (r = -0.52) and with recorded customer satisfaction (r = 0.62); p < 0.001 for both.

Leadership and management without agentic experience

Vendor leadership teams do not yet hold the experience their firms are selling. Senior respondents rated the overall effectiveness of their employer’s top leadership team at 4.09, against 3.60 across all 850 respondents; 39.89% of the senior ratings on that condition fall in the distressed band. Leadership effectiveness tracks outcomes at the engagement level, correlating at r = -0.51 with cancellation and r = 0.61 with recorded customer satisfaction (p < 0.001 for both). Engagements where senior-leadership reporting was 3 or below were canceled in 64.29% of cases (36 of 56), with recorded satisfaction of 25.75%; where it sat at 6 or higher, cancellation fell to 8.57% (3 of 35), with satisfaction of 67.60%.

Leadership’s understanding of agentic AI widens the gap further. Engagements where leadership’s agentic AI knowledge scored 3 or below canceled at 89.66% (26 of 29) and recorded customer satisfaction of 13.66%. Engagements with a leadership agentic AI rating of 6 or higher recorded no cancellations across all 41 engagements and a mean satisfaction of 71.68%.

The pattern extends to mid-tier management. All 7 direct-manager measures separate canceled from active engagements more sharply than any of the 20 employer-capability conditions. The 5 manager-competence measures correlate with cancellation between r = -0.62 and -0.67 and with recorded satisfaction from r = 0.74 to 0.78, p < 0.001 throughout; the record does not settle the direction of cause, and failing engagements may sour their teams on the manager.

The managers running these engagements reported their own lack of experience with agentic AI. 53.33% of the study’s 150 Engagement Managers reported 0 agentic AI experience, with a mean of 0.47 years; no Engagement Manager reported more than 1 year. Among the 31 Portfolio Directors, 64.52% reported 0 experience, with a mean of 0.33 years. Manager inexperience carries a measurable price: engagements whose Engagement Manager reported 0 agentic years canceled at 45.00% (36 of 80), with recorded satisfaction of 38.44%; engagements whose Engagement Manager reported a single year canceled at 28.57% (20 of 70), with satisfaction of 50.04% — a difference significant at p = 0.04.

The missing vendor-side playbook

Vendors lack a mature agentic AI playbook to govern internal operations and the full customer delivery lifecycle. Respondents rated the extent to which internal management processes have been redesigned for agentic delivery at 3.66, with 48.56% distressed — management running an agentic business through a pre-agentic internal process, with no firm-level rules to inherit.

Respondents rated their employer’s ability to manage agentic engagements end-to-end at 3.69, and rated the maturity of the playbook for taking agentic use cases from concept to governed production at 3.65, with 48.57% of ratings in the distressed range. On canceled engagements that reading fell to 2.40; engagements in production stood at 3.94. The absence of a mature playbook led to damaging issue selections in 111 of the 150 engagements. Playbook maturity correlates with cancellation at r = -0.59 and with recorded customer satisfaction at r = 0.71 (p < 0.001 for both).

Vendors are not maturing their customer-facing playbooks fast enough. Among the 94 active engagements, playbook maturity is flat against engagement age: 4.47 up to 3 months old, 4.34 at 4 to 6 months, and 4.45 beyond 7 months. Respondents noted that without an operationally mature client-facing playbook, they invent the rules themselves — junior staff making delivery decisions without formal guidance from their employer, without direct customer involvement, and often outside the visibility of either party’s management team.

The vendor’s own platforms predate agents

Over the last decade, outsourcing vendors have expanded hybrid outsourcing — bundling cloud, SaaS, and other platforms with the human delivery layer — and respondents described these platforms as problematic in an agentic AI context. Respondents rated how well their employer’s bundled legacy platforms are optimized for agentic delivery at 3.70, with 48.36% of ratings in the distressed range; on canceled engagements the reading fell to 2.52. Respondents selected vendor-built platforms that are not agent-ready as the single most damaging issue on 26 engagements.

No agentic talent exists for the vendor to hire

Respondents rated their employer’s supply of qualified agentic AI talent against realized demand at 3.69, with 47.63% of ratings in the distressed range. The rating does not describe a hiring backlog — there is no pool of experienced practitioners for any vendor to recruit, and no vendor in the study had solved the agentic AI talent shortage.

The shortage reached the engagements directly. Respondents selected talent shortage as one of the 5 most damaging issues in 69 of the 150 engagements, and selections concentrated in failure: 57.14% of canceled engagements had a talent-shortage selection, against 39.36% of active ones. Canceled engagements rated the talent supply at 2.46, against 3.93 in production.

With no talent to hire, the vendor’s only remaining lever is training, and respondents rated that as problematic too. Employer training programs for practical agentic AI skills rated 3.69, and for financial-industry knowledge 3.72; on canceled engagements the two readings fell to 2.50 and 2.57. Of the 800 respondents who named a primary source of agentic AI knowledge, only 125 named their employer’s training program (15.62% share) — last among the 6 sources measured, behind vendor and partner training (140), online courses and on-the-job learning (136 each), university coursework (132), and self-teaching (131).

“I’m a fresher and know more about AI than anyone in our executive team.”Individual Contributor, KPO · Active engagement
2.3

Co-management.

Lowest-rated co-management conditionsMean (1–10)Rated distressed
Readiness of the joint management model for agent autonomy3.6050.37%
Quality measures capture the errors that matter3.6151.14%
Clarity on what agents may do without human approval3.6150.25%
Speed to jointly stop an agent behaving unexpectedly3.6250.50%
Agent updates are visible to the other firm before they take effect3.6249.07%
The governance model fits the work performed by agents3.6349.19%
Single shared view of the risks agents introduce3.6448.30%
Human review points placed where errors matter most3.6449.31%
Genuine scrutiny of agent output at review points3.6549.50%
Joint issue resolution matches agent speed3.6549.44%

Table 10 · Lowest-rated co-management conditions

Joint daily management of an engagement whose work is executed by autonomous systems has no precedent in outsourcing practice, and the methods and maturity of joint customer-vendor delivery have yet to be developed in an agentic AI context. Traditional outsourcing governance was paced to people — review meetings, issue resolution, and change boards assume work at human speed. Agentic delivery moves faster than any of those processes and arrives in volumes they cannot clear, and neither the client institutions nor the vendors in this study have built replacements; the market is too young to have produced best practices to adopt.

Among production-stage engagements, 3 conditions rated lowest. Cross-firm visibility of agent actions was rated 3.25, the lowest among the production engagement conditions. Joint human review capacity rated 3.26 — a review built for human throughput cannot clear machine throughput. Joint change control rated 3.3: agents change behavior between formal releases as a model updates, a prompt changes, or upstream data shifts, and respondents frequently make material changes to ensure KPIs are met; a change board built for quarterly release cycles governs none of it, and respondents reported changes reaching production without any joint review.

No co-management playbook exists to adopt. Engagements that survive are those in which the client and vendor build one together before the contract is signed, and the highest-rated engagements keep refining it throughout the engagement.

Co-management becomes critical as the engagement matures

Early-stage engagements rated client-readiness conditions as their weakest category. Among production engagements and all canceled engagements, the 15 weakest-rated of the 100 conditions are all co-management conditions — client and vendor problems are worked down as an engagement matures, and the co-management problem becomes the most impactful category. The 10 lowest-rated co-management conditions range from 3.60 to 3.65 across the study, a spread of 0.5; limited to production engagements, the 3 weakest co-management conditions were rated between 3.25 and 3.27.

Decision rights and accountability

The weakest condition among canceled engagements is clarity on what agents may do without human approval, rated 2.16. Decision rights over agent autonomy are the first question a joint agentic operation must answer, and no canceled engagement had settled it. Severity rises as the rating moves down the management chain toward the daily work: Individual Contributors rated co-management at 3.52 against 4.98 among Portfolio Directors on the same engagements, a gap of 1.46 points.

Cross-firm visibility of agent actions

Agent visibility across the firm boundary was rated 3.25 among production engagements. Asked to name the single most damaging issue in their engagement, respondents chose the agent black box across the boundary 54 times. Visibility of the other party’s agent updates was rated 3.29. Canceled engagements rated the same condition at 2.3; genuine scrutiny of agent output rated 2.16 on canceled engagements, and the quality measures both firms used to validate the other party’s agents rated 2.21. Production engagements rated cross-firm visibility at 3.9, 1.6 points above canceled engagements.

Human review at machine output speed

Joint human review capacity was rated 3.26 across production engagements and led to damaging issue selections in 63 engagements. Placement of review points where errors matter most was rated 3.30, and genuine scrutiny of work product at agentic AI output scale rated 3.4. The study indicates that much of the human review process still operates in a pre-AI context; given non-deterministic AI outputs, humans need to review the work, but the review process must be optimized and scaled specifically for agentic AI.

Joint change control at machine speed

Respondents noted that they must constantly effect changes in response to variations in inputs, customer demands, 3rd-party systems, and the non-deterministic output of AI; the volume of changes teams make exceeds the capability of the customer-vendor change control process, against high demands to achieve KPIs.

Joint change control rated 3.3 among production engagements. The speed mismatch beneath it was measured directly: joint change control processes kept pace with the speed and volume of change at a rating of 3.4 in production, against 2.26 for canceled work. Canceled engagements rated joint governance of human- or agent-made changes at 2.19, against 3.32 in production.

“We were so busy fixing pre-prod issues that we never thought about how to do the actual work.”Engagement Manager, KPO · Active engagement
2.4

The technology ranks behind the institutions.

Lowest-rated technology conditionsMean (1–10)Rated as distressed
Production readiness of the tools for bank-grade work4.5432.04%
Safe handling of sensitive data across outputs and logs4.5432.38%
End-to-end reliability on multi-step tasks4.5532.48%
Accuracy of citations and references against real sources4.5533.50%
Agents took only intended and authorized actions4.5632.07%
Support for reliable pre-release testing4.5632.70%
Resistance to hostile instructions in processed content4.5632.14%
Integrity of persistent memory against contamination4.5731.89%
Ability to explain and provide evidence outputs at audit depth4.5732.75%
Stability of behavior when nothing changed4.5733.00%

Table 11 · Lowest-rated technology conditions

Inside these engagements, agentic AI technology does not meet the standards the financial industry demands. A third of all technology ratings sit in the distressed band, and output-quality defects — including fabricated and inaccurate content — appear throughout the written testimony as a routine operating burden. For institutions whose tolerance for error in regulated processes approaches 0, a technology-layer rating of 4.54–4.66 falls well short of the bank-grade standard. The challenges with the core technology are widely documented beyond this study.

Non-repeating AI output and hallucinations

Agentic AI output is non-deterministic: the same agent, given identical input and configuration, returns different outputs across runs. Respondents rated the tools’ consistency under identical inputs at 4.58. No vendor in the study carried a proven method for evaluating non-repeating output; employer QA capability for such output rated 3.68, with 47.48% of ratings distressed. QA and evaluation gaps led to damaging issue selections in 111 of the 150 engagements, matching the breadth of the missing delivery playbook — an engagement without a playbook improvises its delivery, and an engagement without an evaluation method cannot prove what the improvisation produced.

Teams re-run outputs and compare versions by hand, maintaining non-AI systems in parallel with the agents they operate to validate output and certify samples. None of that labor appears in any client-facing measure. The effort is the cost of operating a non-deterministic system under a deterministic quality model, and the data indicate the vendor absorbs it silently.

“AI invented a husband for a widow on her loan application. I looked at the processing logs. I don’t know why it thought this was important, or how it created this new person.”Portfolio Director, BPO · Active engagement
2.5

The third-party ecosystem ranks last.

Lowest-rated 3rd-party conditionsMean (1–10)Rated distressed
Deliberate handoff of the quiet human fixes for 3rd-party quirks4.5833.54%
Change notice reaching the people running the agents4.5831.95%
Managing dependence on a small number of providers4.5932.04%
Rate limits and quotas under agent volume4.6032.92%
Licenses covering work done by agents rather than named users4.6131.91%
Knowing where humans quietly fixed 3rd-party issues4.6133.08%
Agent dependence on gated external rails4.6130.86%
Fallback readiness for a critical 3rd-party failure4.6132.04%

Table 12 · Lowest-rated 3rd-party conditions

Respondents rated the 3rd-party ecosystem the least severe of the 5 problem areas. Failures inside it remain material and reach every engagement.

Change and service disruption notifications

Change arrives from third parties unannounced — a model version retires, or an API shifts underneath running agents. Upstream data structures change on a cadence the providers alone control, and they owe the engagement nothing. Change notifications and service interruptions are not reaching the delivery team: notice quality rated 4.6, and the paths that would carry notice to the agents’ operators rated the same.

Programmatic access

Respondents rated the availability of effective machine interfaces at 4.63, with 29.80% of ratings distressed. A machine interface replaces a screen built for people. Where APIs exist, their stability was rated 4.66, and the quality of the data arriving through them was rated 4.64. Access to gated external rails, including payment networks and regulatory submission systems, was rated 4.61.

Production engagements rated 3rd-party machine-interface availability at 5.04; canceled engagements rated it 3.14, among the widest spreads in the 3rd-party area. The exposure sits in the ecosystem itself — systems built for human operators now run under agents that cannot use them. Survival and programmatic access move together in this record, whichever drives the other.

“No API, no way in.”Individual Contributor, CXM · Canceled engagement
2.6

The five problem areas across the lifecycle.

An engagement opens against a client institution that has not prepared adequately for agentic AI and has not settled ownership. Delivery then falls to a vendor with at most 2 years of experience doing the work, depending on a workforce without any meaningful agentic AI experience. An engagement that survives those 2 conditions long enough to mature meets the third: no one in this study has built an effective joint co-management discipline. Underneath sits an immature technology, and around it an ecosystem that has not yet evolved for agentic use.

Failure in this study set is institutional in origin rather than technical. Those conditions sit within the buying institution’s authority to change. Part 5 defines the resolution actions.

Continue the report

All parts →
Overview

Report home

The hub: findings, agenda, definitions, and report information.

Read →
Summary

Executive Summary

The sample, the four findings, and the six executive actions.

Read →
Part 1

The Market

Adoption is broad and young, and the losses are already realized.

Read →
Part 3

The Integrity of Performance Information

Reporting the producers rate more likely manipulated than accurate.

Read →
Part 4

The Delivery Workforce

The oversight layer is junior, dissatisfied, and leaving.

Read →
Part 5

Conclusions & the Executive Agenda

Six actions, all within the institution’s own authority.

Read →
Method

Study Information & AI Disclosure

Design, independence, statistics, limits, and the AI-assistance disclosure in full.

Read →