Failing Legitimacy: How AI Can Restore American's Trust in Institutions

Why American health insurance is the highest-leverage test laboratory for smooth and augmentative, rather than human replacement, AI deployment in the economy

πŸ‘€ Aaron Kushner
πŸ“… 2026 · 🏷 AI Deployment · Health Insurance · The Corporate Pivot
Developed with: Claude (Anthropic)
The author declares no competing financial interest or funding from any commercial AI tool provider.

Introduction

A bargain between a private for-profit industry that serves a critical societal need and it's customers, the general public, can fracture without anyone in the industry quite noticing. The legitimacy of American health insurance is one of the most clear live examples of this phenomenon of corporate failure to "read the room". The instinct is to read the fracture as a story about who deserves blame. This paper proposes a structurally different frame. The industry's product is about managing friction between the priorities of profit and public health. Managing the internal corporate friction of such large companies, with overlapping and often contradictory interests entrenched within the rigid corporate structure, takes a dramatic toll on efficiency, forcing productive value to be spent in internal operations and human resources, that could otherwise be available to serve the external customer.  The product is widely accepted to be flawed and exploitative by it's customers, who don't have a choice of another industrial-scale provider class: Socialized medicine is not the answer, it just shifts the corporate bureaucracy over to the purview of the administrative state.

AI offers a path to restore the legitimacy of the entire industry, in the process showing that for this tech revolution, it behooves industry to look at it as a route to providing better value at the same price, rather than simply as a way to cut expensive humans and replace them with machines. Artificial intelligence is the specific tool that can be used to refactor or reform the internal corporate machinery to smooth out the friction surfaces and interfaces, keeping most of the human infrastructure intact, happier, and better performing.

There are two modes of AI deployment inside a corporation that do opposite things. Mode 1[1] takes a human-performed task and removes the human, output cheaper; it is the mode driving the disruption people fear. Mode 2 takes the existing human and strips the paperwork, coordination, and decision overhead around them, redirecting freed time to the parts that actually produced value β€” judgment, navigation, relationship, service. Health insurance is the highest-leverage Mode 2 test laboratory in the American economy. This is a critical decision with real consequences for societal stability, where the current state of affairs in the republic, and therefore Pax of Americana, is much more brittle and on edge than it was at the time of the last major technology disruption. The customers have no control: will the provider choose Mode 1 and simply use AI to cut costs and provide the same problematic service, with all the savings earmarked for short term growth signaling in earnings reports, rather than increasing the value delivered to the customer?

This paper argues that a public commitment to Mode 2 deployment would provide short term relief on the legitimacy issue that is one of the strongest indicators of American decline. It sets out the methodology of an embedded refactor within the health insurance corporate context, and names four concrete refactor targets the diagnosis points at. If it can work for the US healthcare system, it can work at everyone of our institutions, all of which suffer rising legitimacy issues, whether it's the university system and granting agencies, or the whole government system that manages the administrative state. Corporate foundations at corporate behemoths were built two industrial/technological revolutions ago, and have struggled to adapt within the constraints and expedients of modern capitalism in a way that sustainably and effectively serves their customers. The AI revolution is not the next challenge to a creaking system, but the upgrade key, that realigns the compact between corporations and their customers in a way that satisfies the needs of both.

Β§1. The Two Modes

A distinction that dissolves an apparent paradox

Almost every public conversation about AI and work treats AI as one thing, and asks what it will do to one labor market. The framing is wrong at the root. There are two distinct modes of AI deployment inside a corporate environment, and they have opposite vectors. The mode-collapse failure β€” treating humans and machines as the same tool used twice β€” is the failure that drives most of the nihilism about AI in the workforce.

Mode 1: Wholesale replacement

Take a discrete task currently performed by a human. Train or deploy a model that performs the task at acceptable quality. Remove the human from the loop. Book the headcount reduction. This is the mode that sells easily to a CFO on a quarterly call. It is also the mode that produces the visible labor disruption that drives the present anxiety.

Mode 2: Friction removal and redirect

Keep the existing human in the existing seat. Identify the paperwork, the routing, the coordination, the lookup, the form-rephrasing, the recapitulation of information already in the system β€” the activity surrounding the human that does not produce value but consumes most of the time. Refactor that surrounding work with AI. Redirect the freed time to the part of the seat that did produce value β€” judgment, navigation, relationship, service. Keep the headcount. Keep the cost structure. Improve every output across customer, employee, and shareholder. This mode has been named in the augmentation-over-automation literature β€” Brynjolfsson's Turing Trap (2022) is its canonical modern statement, and the empirical evidence that augmentation works in practice is laid out in Brynjolfsson, Li & Raymond's Generative AI at Work (2023). The contribution of this paper is not the naming. It is the substrate-law grounding for why Mode 2 is the architecturally correct deployment, and the methodology for landing it inside a specific industry where the conditions for the test are unusually well co-located.

The two modes use the same underlying instrument. They are not the same deployment. The disruptive mode wins by default because it is the one that maps cleanly onto a quarterly earnings narrative. The refactor mode requires something the replacement mode does not β€” a person on the inside with both vantage and temperament to drive it. That requirement is what makes Mode 2 the harder mode to scale and is also the reason it has been underdeployed.

Naming the distinction is the move that resolves the apparent paradox at the heart of every "AI causes the problem; AI solves the problem" argument. It is only paradoxical if the two modes are collapsed. Mode 1 destroys the seat. Mode 2 refactors the seat. Both are AI. The replacement frame is what should be resisted by policy and by corporate strategy. The refactor frame is what can be deployed at scale starting now, in industries where the conditions favor it. The rest of this paper is about the industry where the conditions are most favorable.

Β§2. Why Health Insurance is a Good Test Laboratory

A $1.4 trillion friction engine, ready to be refactored

Decisions, not things. The American health-insurance industry processes roughly $1.4 trillion in premium revenue annually and sits atop a $4.5 trillion healthcare-spending stack. It does not manufacture a single physical good. Its entire output is decisions about who pays for what: prior authorization decisions, eligibility decisions, network-coverage decisions, denial decisions, appeal decisions, claims-adjudication decisions, formulary decisions. The whole industry is a friction engine that produces decisions and routes paperwork. That is exactly the production model Mode 2 is shaped to refactor β€” decisions live in language and structured data, and the overhead surrounding the decision-maker (form preparation, lookup, routing, escalation, rephrasing) is exactly the surface AI is good at.

Reciprocally hated friction. The American Medical Association's annual prior-authorization physician surveys document a remarkably consistent pattern. The overwhelming majority of physicians report that prior-authorization processes cause care delays. A large share report that the delays have led to abandonment of treatment by patients. A meaningful fraction report that prior authorization contributed to a serious adverse event for a patient in their care. Physicians spend something on the order of a full business day per week, per practice, on PA paperwork that does not produce care. The member side of the same transaction tells the inverse story β€” long phone holds, repeated information requests, denial letters written in language designed for the legal record rather than for human comprehension. The employee side β€” the customer-service representative, the utilization-management nurse, the prior-auth reviewer β€” does not enjoy the work either. It is one of the rare industries where every actor at every position in the value chain finds the work miserable and would actively cooperate with a refactor. That is a rare political condition. It is the condition that makes the test possible.[2]

Regulatory exposure that raises Mode 1's cost. HIPAA, ERISA, state Departments of Insurance, and the new wave of CMS prior-authorization rules all converge on the same baseline: regulated medical-coverage decisions require a human in the loop, with a credentialed clinical or licensed-adjuster role accountable for the final call. This raises β€” but does not preclude β€” Mode 1 at the decision layer. Mode 1 at the decision layer β€” replacing the human reviewer with an automated denial β€” is legally exposed; Mode 1 around the decision layer (intake, drafting, routing, eligibility checking) is permitted, and is in fact what most current industry deployments do. UnitedHealth's nH Predict prior-authorization system is the documented counterexample to a stronger reading β€” Mode 1 at the decision layer with a nominal human reviewer rubber-stamping outputs, producing the 2023 Livingston v. UnitedHealth class action and the October 2024 Senate Permanent Subcommittee on Investigations report. The regulatory shape doesn't mandate Mode 2; it raises the cost of running Mode 1 cleanly enough to make Mode 2 the lower-risk path for any payer whose general counsel is paying attention.[3]

Margins cyclically thin enough to add political cost. Major US health insurers operate at thin and cyclically falling net margins. The NAIC's 2024 Annual Industry Report puts the U.S. health insurance industry's profit margin at 0.8% in 2024, down from 2.2% in 2023, with a net underwriting loss of approximately $1.3 billion across the industry offset only by investment income. In years where the cycle bottoms, the political cost of brittlizing service delivery in pursuit of headcount cuts is hard for any payer to absorb cleanly; in years where margins recover, the cost falls. The cyclical-thinness claim is narrower than "margins preclude Mode 1" β€” UnitedHealth has demonstrably had the margin to run nH Predict despite the cycle. The arithmetic does support Mode 2 β€” keep cost structure flat, dramatically improve unit-output quality, expand margin slowly through reduced rework and reduced appeal volume and reduced member churn. The financial story Mode 2 has to tell is not the headcount-cut story. It is the rework-reduction-and-customer-retention story. That is a story the quarterly call can be taught to tell.[4]

An Example Mode 2 Contract

A Mode 2 commitment is a binding contract clause: when AI tooling absorbs the friction work surrounding an employee's role, the employee retains salary at the prior level for a defined transition window β€” 24 months is a reasonable starting frame. During the window the employee partners with an embedded AI workpartner β€” software that holds the employee's context, conversation history, and the firm's operational state β€” to identify new value-add inside the company. If the first-pass refactor does not surface a fit, separation is mutual and the package is generous: enhanced severance, retraining stipend, references from named senior leadership. The contractual commitment is what differentiates a public Mode 2 deployment from the seat-level four targets in Β§4 β€” those describe what the AI does. This describes what the company commits to do about the human.

Β§3. The Friction Dump Methodology

Embedded research with an AI partner, starting from a grievance log

The methodology this paper proposes for running the experiment is paired with an AI solutions B2B startup, a consulting engagement that could be mutually beneficial, and is free to the corporation. The onboarding/intake process itself turns into something more like a research collaboration between an inside operator and an AI partner. The starting input is a particular kind of artifact: an unstructured, free-associated grievance log produced over roughly a week by a single mid-level insider at a single payer organization. Call it the Friction Dump.

The FD is intentionally low-structure. Free format, the machine handles everything. The instruction to the insider is short. Write down what frustrates, drags, or feels wrong from inside your day. Be granular. Be anecdotal. Interpersonal material is welcome β€” who pushes back on what, whose name comes up in every escalation. If a thought interrupts another thought, write the interruption down and come back, or do not come back. Order matters, including the order things come up unprompted. The stop condition is felt, not scheduled: when there is really nothing left, the insider says so.

The results are completely anonymous and anonymized where viewed by the corporation, a condition that should help encourage the delivery of a rich seat environment surface, enough data to identify the types of friction. Then a solution is formulated collaboratively, and a tool is built by the AI solutions company that completely dissolves the friction point.

The shape of the intake tool is doing specific cognitive work. It is recording three signals an intake form cannot record. The first is order of unprompted surface: what comes up first when no one is asking, the second time, the tenth time. The second is frequency of recurrence: which complaints reappear across days under different surface descriptions. The third is what gets left dangling: the thought that the insider started and did not finish, the workflow whose description ran out before the workflow itself did. All three are visible only in unstructured free-association across time. None of them survive being collapsed into a structured intake form. A consultant cannot extract them. An employee survey cannot extract them. They are produced by the act of... dumping.

The AI's job on the receiving end is also something the existing consulting model cannot do at all. It is to absorb a week of unstructured text, hold it in working memory simultaneously, and surface the structural patterns the insider cannot see from inside the building β€” because the insider is too embedded in the day-to-day to register the recurrence, and because the patterns live in the relations between items the insider already filed under separate mental categories. An AI partner can hold every entry against every other entry, find the recurrences across surface descriptions, identify the dangles, and produce a synthesized map of the operational schema's failure modes with the leverage points marked. The output of the map is then input to a refactor design pass β€” which refactor targets the insider can actually move from where they sit, which require political cover from above, which require the partnership of a sister function.

The methodology generalizes. Any institution in which a thoughtful insider has the vantage but not the integration capacity is a candidate for the same protocol. The reason it works in health insurance first is that the insider population β€” mid-managers in payer organizations, clinicians-in-payer roles, utilization-management nurses with operational scope β€” is unusually large, unusually well-positioned, and unusually well-motivated to participate. The industry has been producing thoughtful insiders frustrated with their own workflows for years. The dump methodology gives them a use for the frustration.

Β§4. The Four Seat-Level Targets (already in industry deployment)

What the industry is doing now β€” and where this paper's contribution is NOT

Seat-level Mode 2 deployment is already underway β€” Anthropic's Claude for Healthcare, OpenAI's ChatGPT Health, Optum's Digital Auth Complete. This paper's distinctive claim is that seat-level deployment alone is insufficient. The institutional-altitude refactor β€” coordinated AI-maintained shared case-state across pillars, policy propagation, claims-to-appeals feedback loops β€” is where Mode 2 either becomes coherent or collapses back into expensive Mode 1. The UnitedHealth nH Predict pattern is what collapse looks like in practice: seat-level Mode 1 with a token human reviewer, no institutional-altitude coherence to catch the systemic drift. The four seat-level targets named below are necessary; they are not sufficient.

  1. Prior authorization as a coherence layer, not a removal layer. AI-prepared dossier (clinical evidence summarized, formulary criteria mapped against the patient's record, precedents from the payer's own past decisions surfaced, draft rationale prepared) presented to the human reviewer β€” the prior-auth nurse, the utilization-management nurse, the clinical pharmacist. The decision stays human. The paperwork around the decision is refactored.
  2. Appeals as AI-drafted and human-checked. AI-prepared synthesis (denial rationale, appeal submission, clinical evidence, policy text and case precedent, draft reasoning pass) presented to the appeals reviewer. The reviewer keeps the decision; throughput rises; the response engages substance instead of templated boilerplate.
  3. Member-service representative tooling β€” situation-aware, not script-driven. AI pre-loads the member's situation (recent claims, pending authorizations, eligibility, the last three interactions across channels, provider network in their geography, formulary status) in the half-second before the rep picks up. The rep, no longer reading from a script and no longer looking up codes, is a person on a call with another person.
  4. Care navigation as proactive matching, not reactive lookup. AI holds the member's clinical history, payer-side provider quality data, geographic and language and scheduling constraints, and actual network availability. The function becomes proactive matchmaking with specific high-fit options β€” not directory recitation.

These four targets are already being deployed by Anthropic's Claude for Healthcare (Banner Health, Novo Nordisk, Sanofi, AdventHealth, Boston Children's, Cedars-Sinai, HCA, Memorial Sloan Kettering, Stanford Medicine, and others; announced January 11, 2026), by OpenAI's ChatGPT Health (announced January 7, 2026), by Optum's Digital Auth Complete (early 2026), and by UnitedHealth's PreCheck workflow.[5] They are well understood, well resourced, and well publicized. The financial press will tell the story under headings like throughput improvement, appeal-cycle compression, call-time reduction, and medical-loss-ratio improvement. This paper does not claim the seat-level four as a novel contribution. The work has been done; the industry is doing it.

What gets less attention β€” and what the rest of this paper turns to β€” is the same Mode-2 refactor one altitude up, inside the institution itself. The seat-level targets are necessary. They are not sufficient. The harder and less-discussed altitude is the one between the seats β€” the spaces between departments where the institution currently fights itself, where state goes incoherent across the boundaries, where the same case can be decided four different ways by four different parts of the same company, and where the deployment patterns the industry is comfortable with at the seat level have not yet been articulated. That is the contribution of this paper. It is the subject of Β§5.

Β§5. The Institutional Altitude

Where the same Mode-2 refactor operates one layer up β€” keeping the institution coherent to itself

Mode-2 at the seat removes the friction surrounding one worker doing one case. Mode-2 at the institutional layer is the same refactor flowing up β€” keeping institutional state coherent across the departments that together produce the member's experience. The two altitudes are not alternatives. They are the same architectural commitment expressed at different scales of the institution, and they hold each other honest. Seat-only deployment produces faster local decisions that contradict each other across departments. Institutional-only deployment produces a coherent corporate posture nobody at any individual seat can actually act on. Both layers running together produce the multidimensional, hierarchically interdependent state architecture that lets the institution behave like a single coherent actor instead of a federation of departments fighting through the member.

A typical major payer is organized into roughly nine functional pillars: product and benefit design; network operations (contracting and provider relations); medical affairs (clinical policy, utilization management, and the medical directorate); claims operations; appeals and grievances; member services; care management; pharmacy; and compliance. Each pillar holds its own state β€” its own policies, its own decision logs, its own member touchpoints, its own information systems. Most of the misery the member experiences is produced not inside a single pillar but at the boundaries between them, where one pillar's state has drifted out of alignment with another's. The institutional-layer refactor is the work of keeping the state aligned. The four targets below are the cleanest cases.

5. Cross-departmental case coherence

A single case β€” a single member, a single condition, a single sequence of decisions β€” currently exists as a different case inside each pillar that touches it. Claims sees one slice. Appeals sees a different slice with different framing. Care management sees a third. Member services sees a fourth. The pillars do not share a coherent representation of the case across boundaries, and so the institution's position on the case is inconsistent at the seams. The institutional refactor is a shared, AI-maintained case-state representation that every touching pillar reads from and writes to. When appeals overturns a denial, the criterion's interpretation in that case updates wherever the criterion lives in any other pillar's decision support; when care management adds a clinical context, the next claims or member-services interaction on the same case sees it; the institution holds one position on the case across the four pillars instead of four positions.

Who keeps the seat
Every pillar that touches a case; no department loses headcount, every department gains shared ground truth
What changes
The institution stops issuing contradictory positions on the same case to different parties
What the press calls it
Operational coherence; rework reduction; reduced grievance escalation; defensible audit posture
An illustrative case β€” before and after

A member with a complex autoimmune condition has a specialty-drug claim denied in March. She appeals; the appeals reviewer overturns the denial in May, citing the documented response to prior therapies under criterion 3.4.b. The drug is dispensed in June. In August the refill claim arrives. Claims denies again β€” because the May appeals overturn never landed in claims's decision support, the reasoning was filed in the appeals system the claims team does not query, and the precedent the appeals reviewer set is invisible to the next claims adjudicator. The member calls member services to ask why a drug she just got approved is being denied again. Member services has no access to the appeals reasoning either and tells her to file another appeal. She files another appeal. Six months of bouncing through four pillars on a case the institution already decided once.

The May overturn lands in shared case state. The criterion's interpretation in this case is now part of the case's living record, visible to every pillar that touches the case. The August refill claim pre-loads the May precedent and auto-authorizes under the same criterion the appeals reviewer already applied. The member services rep, if the member calls anyway, sees the full institutional history and can explain it in one sentence. The institution holds itself accountable to its own prior decisions instead of asking the member to re-prove the case quarterly. Aggregate effect: the rework volume the payer was paying for at $1,200 a Level-2 appeal β€” much of it self-inflicted by state incoherence β€” drops by a measurable fraction.

6. Policy-to-operations propagation

When medical affairs updates clinical policy, the update has to land in every operational pillar that enforces or explains it β€” utilization management, claims operations, network ops, member services, appeals β€” and currently it does not, on any predictable timetable. The lag between policy and operations is the most consequential structural friction inside a payer that does not show up in any single seat's queue. Members get told the old answer for months after the new policy is in force. Reviewers apply old criteria to new cases. Appeals reverse decisions made under the prior version. The institution speaks with five voices about the same policy because the five operational pillars are reading from five different snapshots of when the policy was last refreshed. The institutional refactor is a shared policy state every operational pillar reads from and AI-prepared decision support that re-renders the moment the policy state changes.

Who keeps the seat
Medical affairs, UM, claims, network ops, member services, appeals; no one loses, everyone reads from the same version
What changes
Policy-to-operations lag collapses from quarters to days; the institution speaks with one voice within 48 hours of a policy change
What the press calls it
Operational agility; regulatory-update responsiveness; reduced compliance exposure
An illustrative quarter β€” before and after

Medical affairs updates the clinical policy on a class of GLP-1 medications in early Q1 β€” narrowing the indications under which the drug is covered. The update memo goes to the UM team, which incorporates it into reviewer checklists over six weeks. Network operations learns about it at the next quarterly contract review, three months in. Member services learns about it when members start calling confused, four months in. Appeals reviewers learn about it from a Level-3 case six months in, while still upholding old denials made under the prior policy. Compliance learns about it from a regulator's letter eight months in. The institution has been operating under five different versions of the same policy for two quarters. Member trust takes a measurable hit. Two state DOIs open inquiries.

Medical affairs publishes the update to the shared policy state on a Tuesday morning. By Wednesday morning, the UM reviewer's decision support reflects the new criterion. The contract templates the network team uses pick up the change. The member services situational pre-load includes the new policy as context the moment a member call comes in. The appeals queue is filtered for cases the new policy actually resolves, surfacing them to the appropriate reviewers. The compliance team gets the propagation log as audit evidence. The institution speaks with one voice within 48 hours. The regulators' inquiries do not get filed because the institutional record demonstrates the policy was applied consistently from day one.

7. Claims-to-appeals feedback loop

Appeals exists because claims sometimes makes the wrong call. Currently, when appeals overturns a claims decision, the overturn is a one-off event recorded in the appeals system. It does not feed back into the claims team's criteria interpretation. The same kind of case keeps getting denied at the claims layer and overturned at the appeals layer β€” the institution pays twice for the same decision and the member pays in delay. The institutional refactor is a closed feedback loop: every appeals overturn is treated as quality signal to the underlying claims decision logic, AI-categorized by the criterion-interpretation it represents, and surfaced to the claims team as a calibration update. Claims decisions in the same category begin tracking the appeals layer's accumulated reasoning. The structural inconsistency between the two pillars closes.

Who keeps the seat
Claims adjudicators and appeals reviewers; both seats keep their role, both seats stop fighting through the member
What changes
Recurring case patterns the appeals layer is reversing stop being made at the claims layer in the first place
What the press calls it
Appeal-rate reduction; first-pass-yield improvement; medical-loss-ratio improvement through reduced administrative cost
An illustrative pattern β€” before and after

The appeals team notices, informally, that they are reversing roughly 70% of denials in a particular code category β€” physical therapy for post-surgical knee rehabilitation extending beyond eight sessions. The pattern is visible to anyone who works in appeals long enough to see it. It is not visible to anyone in claims, who continues denying the same cases at the same rate. There is no institutional mechanism to translate "we keep reversing these" into "stop denying these." Twice a year someone proposes a process change. Nothing happens. The institution pays for the same decision twice β€” once in the claim denial, once in the appeal review β€” on every case in the pattern. Members miss two to four weeks of care per case in the gap. Providers learn to file every denied case as an appeal as a matter of policy, which inflates the appeals queue further.

Every appeals overturn is categorized in shared institutional state by the criterion-interpretation the reviewer applied. When the post-surgical knee-rehab pattern crosses a quality-signal threshold β€” say, 60% of denials in the category being reversed within 30 days β€” the claims decision support for that category is updated automatically and the medical director's office is notified that a policy-clarification ratification is required. The next denials in the category begin tracking the appeals layer's reasoning. The institution learns from itself. The Level-1 appeal volume in the category drops 70% over the next two quarters because the institution stops generating the denial in the first place. The appeals team's throughput rises further because they are no longer adjudicating the easy ones; they are doing the genuinely hard cases.

8. Member services as the absorption layer made whole

Member services is the institutional pillar that absorbs every other pillar's friction. When the product team designs a benefit operations cannot enforce coherently, member services takes the calls. When medical affairs updates a policy late, member services explains the inconsistency. When claims denies a case appeals will overturn, member services is the first place the member calls. Currently, the rep handling the call has visibility into the slice their CRM shows them and nothing else. The structural reason member services is the most-burnt-out pillar in the company is not that the work is uniquely hard at the seat; it is that the seat is sitting at the bottom of every other pillar's incoherence. The institutional refactor gives member services the same shared case state every other pillar reads from. The rep on the call sees the full institutional history of the case β€” what claims decided, what appeals is reviewing, what care management is coordinating, what the policy update was, where the boundary friction is. The rep stops being the absorption layer for invisible structural conflict and starts being the visible coordinator across the pillars.

Who keeps the seat
Member services reps, escalation specialists, complex-case advocates; the seat is the same, the vantage is whole-institution instead of single-pillar
What changes
The rep becomes the coordinator across the pillars instead of the absorber of inter-pillar friction
What the press calls it
Member-services attrition reduction; first-call-resolution lift; net-promoter-score recovery; retention improvement
An illustrative day β€” before and after

A member with a chronic condition calls member services. She has called member services twice in the past four months. The first call was about a billing question from a denied claim that was then overturned on appeal. The second was about a confusing letter she got after medical affairs updated a coverage policy. She is now calling because she got assigned to a care manager who scheduled her for a service her health plan does not cover. The rep handling the current call sees none of the prior context β€” not the appeal, not the policy update, not the care manager's scheduling. The rep takes the member's question at face value, calls the care management team to verify the service, gets routed to a voicemail, asks the member to hold, gets back on the line with information that contradicts what the care manager told the member two weeks ago. The member is now angry at the institution as one entity even though no single person in the institution did anything wrong. The rep ends the call burned out from absorbing the structural conflict.

Same call. The rep's screen pre-loads the member's full institutional history across pillars: the appeals overturn from four months ago and its reasoning, the policy update that triggered the second call, the care plan the care manager put together two weeks ago including the scheduled service. The rep can see immediately that the scheduled service falls outside the policy update's narrowed criteria β€” the care manager scheduled it under the pre-update interpretation and has not been re-prompted. The rep can either resolve it on the call (explain the situation, route the case back to care management with the policy context attached, offer an alternative covered service) or coordinate across the pillars without making the member tell the story again. The rep ends the call having coordinated, not having absorbed. The structural friction the rep was carrying is visible now, and visible structural friction can be designed against.

The four institutional targets reinforce the four seat-level targets in a specific way that is worth naming. Each seat-level refactor depends on institutional state coherence to do its work properly. Diane's prior-auth dossier has to reflect the policy update medical affairs made yesterday, not last quarter. Marcus's appeals reasoning has to feed back into claims's decision criteria, not get filed in a separate system the claims team does not search. Aisha's situational pre-load has to include the care plan Jordan put together for the same member three months ago, not stop at the slice the CRM shows her. The institutional layer holds the seat layer honest. Without the institutional layer, the seat refactors run faster against incoherent state, which produces faster local decisions that contradict each other across the institution β€” a worse outcome than the status quo in some scenarios because the speed amplifies the incoherence. With the institutional layer, the institution stops fighting itself, and the seat refactors get to do what the target-grids say they do. The two altitudes are one refactor.

Β§6. Coda β€” Why This Generalizes

The first deployable test case of a broader pattern

This paper has argued that American health insurance is the highest-leverage test laboratory in the economy for the second mode of AI deployment. It is the test laboratory because four conditions co-occur with unusual cleanness β€” decisions as production, reciprocally hated friction, regulatory cover for humans-in-the-loop, and margin structure that precludes the alternative. The seat-level refactor targets are well understood and already being actively deployed by major payers and frontier-AI healthcare partnerships as of 2026; this paper's distinctive contribution is the four institutional-altitude targets that follow from the same diagnosis. The two altitudes are one refactor, not two: the same architectural commitment expressed at different scales of the institution, with each holding the other honest.

The argument does not stop at this industry. Health insurance is the first deployable test case of a pattern that can, with adjustments, generalize across every institution whose product is decisions and whose operating cost is the friction surrounding the decision-makers. That pattern includes commercial banking's compliance and lending operations, legal services' document and intake operations, public-sector benefits administration, large-scale procurement, and substantial portions of utility and telecommunications customer operations. The conditions are not all present in every case. Some industries have the friction but not the regulatory cover. Some have the regulatory cover but not the political will. Some have everything except the inside operators with the temperament to drive the refactor. The conditions matter; they are why this paper is specific about which test laboratory to run first.

The deeper claim is that the disruption-versus-redirect frame is a load-bearing distinction the broader public conversation about AI and work has not yet made. The replacement narrative is the default because it is the easiest one to monetize on a quarterly call and the easiest one to fear. The refactor narrative is the harder one to monetize and the harder one to communicate, and it is the one the substrate-asymmetry argument elsewhere in this series suggests is the only one that is structurally stable. Humans hold the substrate that produces judgment, presence, recognition, and stake β€” the parts of work that cannot be performed by a frozen statistical surface. The architecture that respects the asymmetry keeps the human in the seat and refactors the work around the seat. The architecture that does not respect the asymmetry removes the human and discovers, after the removal, what could not be replaced.

Health insurance is where the experiment can be run now, by the people who are already in the building, against a friction surface that every actor in the system already wants reduced. Run the experiment here. Publish the results. The rest of the economy will read them.

References and source notes

  1. Gibbons, M., Limoges, C., Nowotny, H., Schwartzman, S., Scott, P., & Trow, M. (1994). The New Production of Knowledge: The Dynamics of Science and Research in Contemporary Societies. Sage Publications. The labels "Mode 1" and "Mode 2" in this paper follow the augmentation-vs-automation distinction in the AI-labor literature (Brynjolfsson, 2022). They are unrelated to the "Mode 1 / Mode 2" usage in innovation management theory established by Gibbons et al., which contrasts academic-disciplinary knowledge production (Mode 1) with industry-linked transdisciplinary knowledge production (Mode 2). The label collision is noted here so that management-science readers do not import the wrong framework.
  2. American Medical Association. Prior Authorization Physician Survey (2024-2025 wave, released 2025). 1,000 practicing physicians, December 2024 administration. Documents physician-reported burden (13 hours/week, 39 PAs/week, on average), treatment-abandonment rate (82% of physicians report patients commonly abandoning treatment due to PA), and adverse-event correlation (29% of physicians report PA contributed to a serious adverse event in their care). ama-assn.org/system/files/prior-authorization-survey.pdf. Survey landing page: ama-assn.org/practice-management/prior-authorization/ama-prior-authorization-physician-survey.
  3. Centers for Medicare & Medicaid Services. CMS Interoperability and Prior Authorization Final Rule (CMS-0057-F, 2024). Final rule released January 17, 2024. Establishes regulatory framework for electronic PA, expedited decision timelines (72 hours expedited / 7 days standard), required human clinical review for many cases. The rule explicitly preserves human review at the decision layer; it does not preclude AI around the decision layer. Implementation deadlines: January 1, 2026 for most provisions; January 1, 2027 for API requirements. cms.gov/initiatives/burden-reduction/overview/interoperability/policies-regulations/cms-interoperability-prior-authorization-final-rule-cms-0057-f.
  4. National Association of Insurance Commissioners. U.S. Health Insurance Industry 2024 Annual Results. Industry profit margin 0.8% in 2024 (down from 2.2% in 2023); net underwriting loss ~$1.3B across the industry; remained profitable only via $13.9B investment income. Comprehensive hospital and medical line had $3.1B underwriting profit; Medicaid lost $3.0B; Medicare lost $2.9B. content.naic.org/sites/default/files/2024-annual-health-industry-commentary.pdf. Companion: KFF analysis at kff.org/medicare/health-insurer-financial-performance/.
  5. Industry deployments of seat-level Mode-2 targets as of 2026: Anthropic Claude for Healthcare (January 11, 2026, deployed with Banner Health, Novo Nordisk, Sanofi, AdventHealth, Boston Children's, Cedars-Sinai, HCA, Memorial Sloan Kettering, Stanford Medicine and others); OpenAI ChatGPT Health (January 7, 2026); Optum Digital Auth Complete (early 2026); UnitedHealth PreCheck workflow. These deployments operationalize the seat-level four refactor targets (Β§4) and confirm that the seat-level Mode-2 pattern is well understood and well resourced. They do not address the institutional altitude (Β§5), which is this paper's distinctive contribution.
  6. Brynjolfsson, E. (2022). The Turing Trap: The Promise & Peril of Human-Like Artificial Intelligence. arxiv 2201.04200. The modern progenitor of the substitution-vs-augmentation distinction that Mode 1 / Mode 2 inherits. Brynjolfsson's argument is broadly economic and political (workers lose bargaining power under substitution); this paper supplies the substrate-law structural grounding (staking as substrate property) underneath that argument. arxiv.org/abs/2201.04200.
  7. Brynjolfsson, E., Li, D., & Raymond, L. (2023). Generative AI at Work. NBER Working Paper 31161. The empirical anchor for the Mode 2 deployment pattern: a study of 5,179 customer-support agents shows 14% average productivity gain from generative-AI assistance, 34% for novice workers, minimal for experienced workers β€” the rigorous evidence for what this paper's refactor targets argue narratively. nber.org/papers/w31161.
  8. Acemoglu, D., & Restrepo, P. (2018). The Race Between Man and Machine: Implications of Technology for Growth, Factor Shares, and Employment. American Economic Review 108(6). The economic-growth-modeling case for understanding automation as displacement-plus-reinstatement: automation displaces labor from existing tasks; new tasks reinstate labor demand. The structural mechanism behind why some automation cycles produce job loss and others produce job creation. Adjacent prior art for the Mode 1 / Mode 2 distinction at the macroeconomic layer.
  9. Acemoglu, D., & Restrepo, P. (2019). Automation and New Tasks: How Technology Displaces and Reinstates Labor. Journal of Economic Perspectives 33(2). The empirical-historical follow-up to the 2018 paper, documenting the displacement-reinstatement balance in US labor markets 1947–2017. The conclusion that recent decades show more displacement than reinstatement is the macroeconomic backdrop against which this paper's Mode-2-as-deliberate-commitment argument is positioned.
  10. Autor, D. (2015). Why Are There Still So Many Jobs? The History and Future of Workplace Automation. Journal of Economic Perspectives 29(3). The canonical statement that machines have always complemented human labor where the labor is non-routine, and that the routine-vs-non-routine distinction (rather than skill level per se) is the right cut. Antecedent to the Brynjolfsson augmentation/substitution distinction; this paper's substrate-law-grounded contribution provides the deeper "why" underneath the non-routine category.
  11. Raisch, S., & Krakowski, S. (2021). Artificial Intelligence and Management: The Automation-Augmentation Paradox. Academy of Management Review 46(1). The management-strategy treatment of the automation/augmentation choice as a paradox firms face simultaneously rather than sequentially. This paper's Mode 1 / Mode 2 framing is the substrate-law-grounded resolution of the paradox: at substrate-asymmetric tasks, augmentation is structurally correct; the paradox dissolves when the substrate property is recognized.
  12. Wilson, H.J., & Daugherty, P.R. (2018). Collaborative Intelligence: Humans and AI Are Joining Forces. Harvard Business Review (July–August 2018). The HBR popularization of the augmentation case for business audiences. This paper's contribution past this prior art is the substrate-law structural grounding for why collaborative intelligence is the architecturally correct deployment, not just the higher-performing one empirically.
  13. Centers for Medicare & Medicaid Services, Office of the Actuary. National Health Expenditure Projections. Total US healthcare-spending stack and payer-share breakdown.