Berkeley Strategy Group · for Microsoft · August 2026

The AI Wellness Initiative

Evaluating behavioral, managerial, and organizational readiness for AI-driven workflow reinvention — and the one tool the evidence said was worth building.

21
qualitative sessions
13
hypotheses tested
4
wellness pillars
7+1
recommendations & one build
Why this exists

Usage data tells Microsoft what people click. It doesn't say why they hesitate.

Across six weeks, BSG ran structured interviews with 11 enterprise workers, 7 managers, and 3 early-career employees — every transcript memoized against a controlled, seven-dimension vocabulary — then scored thirteen hypotheses against that evidence and against Microsoft's own 2026 Work Trend Index.

Microsoft has the usage data, but this engagement adds the layer underneath: why people adopt, hesitate, or disengage, and what the organization can do about it. — Project Overview
The evidence base

Fluency is a progression, not a switch — and each step is gated by something different

No participant's friction was about prompting mechanics. Every barrier was trust, discovery, or governance — and none of it shows up on an adoption dashboard.

TIER 1

Assistant

Reviews every output. AI drafts, the worker owns all judgment — prompt-and-review.

Gated by trust: a worker won't hand off a task class until they've watched the tool get it right enough times to predict it.
TIER 2

Delegate / Coworker

Delegates a defined task class after testing it, then spot-checks.

10 of 21 sessions
Gated by discovery: which tool fits which task, how to chain several together, how to recover when a step breaks.
TIER 3

Autopilot

Sets the workflow running; audits the exception, not the output.

7 of 21 sessions
No higher rung to unlock — the risk is over-delegation, or “AI slop,” where output ships without judgment.
What the ladder predicts: the delegation gap is bigger than the prompting gap (H3, strongly supported). T1→T2 requires trust built through reversible, low-stakes testing; T2→T3 requires discovery — knowing which tools to chain and how to recover when they fail. A prompting curriculum moves no one from T1 to T2.
Findings

Four pillars of AI wellness, each grounded in observed behavior

Not an attitude problem — nearly everyone across every cohort is net energized. The bottleneck is structural, and it splits cleanly into four questions.

Security

Do people still feel their role and contribution matter as AI improves?
Anxiety doesn't fall as fluency rises — it relocates from competence to worth.
So whatWorkers need contribution recognized; managers need the expectations set above them calibrated. Two different interventions, one pillar.

Oversight

Can people tell when to trust what AI produces, and when to check it?
Workers delegate what they've watched work, not what they've been trained on.
So whatOffer a sanctioned way to watch the tool act without consequence, then a way to test its validity before release.

Guidance

Do people know what the tools can already do, and what good use looks like?
The ceiling is knowing what “good” looks like — discovery today is accidental.
So whatTwo products, not one: peer-practice visibility for workers, a stated benchmark for managers.

Enablement

Does using AI feel safe, visible, and rewarded by the people around them?
Adoption is set by the room — what's visible, what's safe, what managers reward.
So whatGive managers sanctioned, no-output time to experiment, plus team-level visibility into whether it worked.
Where BSG and Microsoft agree

The internal interviews and the 2026 Work Trend Index describe the same organization

01
The internal and external pictures generally match.Roughly 8 of 13 hypotheses are corroborated by both BSG's interviews and Microsoft's own WTI 2026 data.
02
The bottleneck is organizations, not individuals.Culture, manager support, and governance drive roughly 2× the impact of individual mindset — 67% vs. 32% (WTI 2026).
03
Manager modeling is the highest-leverage, least-built lever.Managers who visibly model AI use drive a 17–30 point lift in trust and critical thinking — yet only 26% of workers see clear leadership alignment.
04
Diverging evidence is a diagnostic flag, not a disproof.“Accountability without visibility” wasn't observed internally, but is empirically real at industry scale (IBM Watson Health, Amazon).
Hypothesis scorecard · H1–H13
Strongly supported (5) Mixed (3) Contradicted (2) Probed, absent (2)
H11 · Fourier = progression signal — strongly supported. “The delegation-threshold crossing is invisible to every tool participants described.”
Synthesis

Seven moves, sequenced by pillar and checked against what Microsoft already ships

Effort and impact are rated per move. Most of the friction here is behavioral — five of seven extend a program that already exists; two are genuine white space.

Security

R1
Re-aim adoption programs at the anxiety object

Programs sold as “reduce anxiety” miss advanced workers, whose worry has moved from competence to standing. Aim the framing at where contribution is recognized.

Extends existing
Effort
Impact

Oversight

R2
Stand up recommendations-only trust sandboxes

A supervised, recommend-only space to watch AI act before handing over control. Delegation trust is built through reversible testing, not training.

Extends existing
Effort
Impact
R3
Name over-delegation in team quality norms

Codify “thought partner, not author.” The failure mode workers name is over-delegation, and no current norm names it.

White space
Effort
Impact

Guidance

R4
Ship a searchable, role-tagged Copilot inventory

A catalog of what Copilot can already do, tagged by role. Workers asked for this by name — it moves them up a level faster than more trust content.

Extends existing
Effort
Impact
R5
Build a cross-team discovery channel

A structured way for teams to share working AI patterns. Today the best uses spread by accident, one hallway conversation at a time.

Extends existing
Effort
Impact

Enablement · highest-leverage

R6
Make manager modeling visible and expected

The single behavior the Work Trend Index ties to a 17-point adoption lift — and the managers BSG observed don't do it. Currently white space: Microsoft has no interface that makes modeling visible.

White space
Effort
Impact
R7
Split the playbook by cohort

A policy-and-access lane for managers, a time-and-pathway lane for ICs. The binding constraint inverts between them — one playbook can't fit both.

White space
Effort
Impact
The tool question

Almost none of this needs software. One gap does.

Most of R1–R7 are norms, manager behaviors, or process changes — software would be the wrong instrument. The evidence pointed to exactly one tool-shaped gap, and it's already built.

Question one

Which evidenced gap does it close? Progression is invisible (H11). Every tool measures usage — “Copilot used X times” — but the event that actually matters, a workflow crossing the delegation threshold, goes unrecorded. One manager in the study couldn't even see his own team: “we're too small.”

Question two

What actually needs a tool? Almost nothing — which is the point. The single tool-shaped gap is progression visibility: a task→tool finder for individual contributors, and a “did the work actually move up” view for managers.

The call: build one thing. Ship the progression-visibility tool now; defer the manager tracking panel until the governance ruling in R4 clears it; drop everything else — most of the rest is behavior, not software.
Microsoft's dashboards measure deployment. Fourier measures whether the work has changed.
LIVE APPLICATION
fourier-dg2.pages.dev
Open Fourier ↗

Fourier is a live Cloudflare app — onboard, add your own tools, and generate a 90-day plan right here.

Open full screen ↗
Built by Kylie Marcisz · Fourier & positioning, BSG engagement team
External check

Failed rollouts outran trust. Scaled ones paired AI with redesign.

Capability outran the systems built to use it confidently
Klarna
Reversed a ~700-role customer-service AI swap after quality suffered — delegation without staged trust-building fails.
IBM Watson Health
Lost hospital contracts after the tool, trained on hypothetical data, was relied on clinically — accountability was granted before visibility existed.
Amazon Hiring AI
Scrapped after it systematically downgraded female candidates — bias found only after live deployment.
ServiceNow
Its own AI maturity index fell 44→35 year over year, even as 90% of IT tickets went AI-handled — broad adoption isn't deep, governed adoption.
AI investment paired with organizational redesign
JPMorgan Chase
LLM Suite grew from 0 to ~200,000 onboarded users in 8 months via a secure sandbox model.
Moderna
750 custom GPTs built within two months, anchored by a 100-person peer-champion program.
EY
EYQ reached 400K employees and 115K monthly users, with 1,000+ field ideas distilled to 8 repeatable patterns.