Berkeley Strategy Group · for Microsoft · August 2026

The AI Wellness Initiative

Evaluating behavioral, managerial, and organizational readiness for AI-driven workflow reinvention — and the one tool the evidence said was worth building.

21
qualitative sessions
13
hypotheses tested
4
wellness pillars
7+1
recommendations & one build
Why this exists

Usage data tells Microsoft what people click. It doesn't say why they hesitate.

Across six weeks, BSG ran structured interviews with 11 enterprise workers, 7 managers, and 3 early-career employees — every transcript memoized against a controlled, seven-dimension vocabulary — then scored thirteen hypotheses against that evidence and against Microsoft's own 2026 Work Trend Index.

Microsoft has the usage data, but this engagement adds the layer underneath: why people adopt, hesitate, or disengage, and what the organization can do about it. — Project Overview
The evidence base

Fluency is a progression, not a switch — and each step is gated by something different

No participant's friction was about prompting mechanics. Every barrier was trust, discovery, or governance — and none of it shows up on an adoption dashboard.

TIER 1

Assistant

Reviews every output. AI drafts, the worker owns all judgment — prompt-and-review.

Gated by trust: a worker won't hand off a task class until they've watched the tool get it right enough times to predict it.
TIER 2

Delegate / Coworker

Delegates a defined task class after testing it, then spot-checks.

10 of 21 sessions
Gated by discovery: which tool fits which task, how to chain several together, how to recover when a step breaks.
TIER 3

Autopilot

Sets the workflow running; audits the exception, not the output.

7 of 21 sessions
No higher rung to unlock — the risk is over-delegation, or “AI slop,” where output ships without judgment.
What the ladder predicts: the delegation gap is bigger than the prompting gap (H3, strongly supported). T1→T2 requires trust built through reversible, low-stakes testing; T2→T3 requires discovery — knowing which tools to chain and how to recover when they fail. A prompting curriculum moves no one from T1 to T2.
Comparison spine

Seven dimensions, three cohorts — where the story matches and where it splits

Twenty-one transcripts, scored against a controlled vocabulary. The heat below reads warm where a dimension diverges by cohort, cool where it holds across all three.

Dimension
Enterprise workers
Managers
Early-career
Signal
D1 · Anxiety object
WorkersCompetence
ManagersNone observed / undefined bar
Early-careerFalling-behind
Diverges
D2 · Pressure source
WorkersSelf
ManagersSelf
Early-careerSelf
Aligns
D3 · Error ownership
WorkersFully self
ManagersFully self
Early-careerShared
Generally aligns
D4 · Skill direction
WorkersCompounding
ManagersCompounding
Early-careerPivoting
Aligns
D5 · Least-equipped
WorkersTool/task selection
ManagersFeels equipped
Early-careerTool mechanics
Diverges
D6 · Org friction
the key row
WorkersNo time or training
ManagersGovernance / security walls
Early-careerUnclear guidance
Diverges
D7 · Net stance
WorkersEnergized / pragmatic
ManagersEnergized
Early-careerEnergized
Aligns
The key row is D6. Workers are blocked by time and training; managers are blocked by governance and access. One playbook cannot fix both — it feeds R7's split-by-cohort call directly.
Findings

Four pillars of AI wellness, each grounded in observed behavior

Not an attitude problem — nearly everyone across every cohort is net energized. The bottleneck is structural, and it splits cleanly into four questions.

Security

Do people still feel their role and contribution matter as AI improves?
Anxiety doesn't fall as fluency rises — it relocates from competence to worth.
So whatWorkers need contribution recognized; managers need the expectations set above them calibrated. Two different interventions, one pillar.

Oversight

Can people tell when to trust what AI produces, and when to check it?
Workers delegate what they've watched work, not what they've been trained on.
So whatOffer a sanctioned way to watch the tool act without consequence, then a way to test its validity before release.

Guidance

Do people know what the tools can already do, and what good use looks like?
The ceiling is knowing what “good” looks like — discovery today is accidental.
So whatTwo products, not one: peer-practice visibility for workers, a stated benchmark for managers.

Enablement

Does using AI feel safe, visible, and rewarded by the people around them?
Adoption is set by the room — what's visible, what's safe, what managers reward.
So whatGive managers sanctioned, no-output time to experiment, plus team-level visibility into whether it worked.
Hypothesis scorecard

Thirteen hypotheses, scored against every transcript

A probed null is a finding too. Five hold up strongly, three read mixed, two are contradicted outright — and H11 is the one that built Fourier.

H1
Anxiety transforms by tier
Strongly supported
H2
Prompt literacy fading
Mixed
H3
Delegation > prompting gap
Strongly supported
H4
Manager reinvention gap
Strongly supported
H5
Seniority inverts at Autopilot
Contradicted
H6
Shadow AI = tier mismatch
Contradicted
H7
Accountability w/o visibility
Probed, absent
H8
MSFT 6–12 months ahead
Supported, complicated
H9
Early adopters plateau
Probed, absent
H10
Trust calibration norm
Strongly supported
H11
Fourier = progression signal
Strongly supported
H12
Reinvention harder for managers
Mixed
H13
Apprenticeship inverting
Mixed
H11, the load-bearing one: “The delegation-threshold crossing is invisible to every tool participants described.” That gap is what Fourier was built to close — see The Tool Question below.
01
The internal and external pictures generally match.Roughly 8 of 13 hypotheses are corroborated by both BSG's interviews and Microsoft's own WTI 2026 data.
02
The bottleneck is organizations, not individuals.Culture, manager support, and governance drive roughly 2× the impact of individual mindset — 67% vs. 32% (WTI 2026).
03
Manager modeling is the highest-leverage, least-built lever.Managers who visibly model AI use drive a 17–30 point lift in trust and critical thinking — yet only 26% of workers see clear leadership alignment.
04
Diverging evidence is a diagnostic flag, not a disproof.“Accountability without visibility” wasn't observed internally, but is empirically real at industry scale (IBM Watson Health, Amazon).
Manager capability model

Encouraging AI use isn't the same as redesigning work around it

Six capabilities, observed in practice across seven manager transcripts. The redesign capabilities — not the encouragement ones — are the ones missing.

0 / 7
have had the “What is your job now?” conversation
3 / 7
visibly model their own AI use — the 17-point-lift behavior
4 / 7
have built no enablement infrastructure
1 · Verification standard-setting
Verification built into the workflow, risk-calibrated by output type.
2 · Psychological safety modeling
Manager models their own AI use openly — the behavior tied to the 17-point lift.
3 · Work redistribution by readiness
Assigns work to AI vs. people by readiness; advances the reluctant, restrains the over-eager.
4 · Workflow scope redesign
Explicitly renames the job: what AI owns, what's newly possible, what the team no longer does.
5 · Mandate translation
Generates direction rather than passing pressure down; turns “go faster” into specific priorities.
6 · Enablement infrastructure designthe one clean differentiator
Self-built infrastructure that runs without the manager present — certified prompts, agents, a review pipeline.
The highest-leverage gap in the study (H4, strongly supported): no manager has an individual-level readiness signal, and none has been asked what their own job becomes.
Synthesis

Seven moves, sequenced by pillar and checked against what Microsoft already ships

Effort and impact are rated per move. Most of the friction here is behavioral — five of seven extend a program that already exists; two are genuine white space.

Quick wins Major bets Fill-ins Reconsider LOW MED HIGH LOW MED HIGH EFFORT → IMPACT → R1 R2 R3 R4 R5 R6 R7
Effort vs. impact per recommendation, as assessed in the readout — the cluster sits high-impact because most of what's missing is behavioral, not technical.

Security

R1
Re-aim adoption programs at the anxiety object

Programs sold as “reduce anxiety” miss advanced workers, whose worry has moved from competence to standing. Aim the framing at where contribution is recognized.

Extends existing
Effort
Impact

Oversight

R2
Stand up recommendations-only trust sandboxes

A supervised, recommend-only space to watch AI act before handing over control. Delegation trust is built through reversible testing, not training.

Extends existing
Effort
Impact
R3
Name over-delegation in team quality norms

Codify “thought partner, not author.” The failure mode workers name is over-delegation, and no current norm names it.

White space
Effort
Impact

Guidance

R4
Ship a searchable, role-tagged Copilot inventory

A catalog of what Copilot can already do, tagged by role. Workers asked for this by name — it moves them up a level faster than more trust content.

Extends existing
Effort
Impact
R5
Build a cross-team discovery channel

A structured way for teams to share working AI patterns. Today the best uses spread by accident, one hallway conversation at a time.

Extends existing
Effort
Impact

Enablement · highest-leverage

R6
Make manager modeling visible and expected

The single behavior the Work Trend Index ties to a 17-point adoption lift — and the managers BSG observed don't do it. Currently white space: Microsoft has no interface that makes modeling visible.

White space
Effort
Impact
R7
Split the playbook by cohort

A policy-and-access lane for managers, a time-and-pathway lane for ICs. The binding constraint inverts between them — one playbook can't fit both.

White space
Effort
Impact
The tool question

Almost none of this needs software. One gap does.

Most of R1–R7 are norms, manager behaviors, or process changes — software would be the wrong instrument. The evidence pointed to exactly one tool-shaped gap, and it's already built.

Question one

Which evidenced gap does it close? Progression is invisible (H11). Every tool measures usage — “Copilot used X times” — but the event that actually matters, a workflow crossing the delegation threshold, goes unrecorded. One manager in the study couldn't even see his own team: “we're too small.”

Question two

What actually needs a tool? Almost nothing — which is the point. The single tool-shaped gap is progression visibility: a task→tool finder for individual contributors, and a “did the work actually move up” view for managers.

The call: build one thing. Ship the progression-visibility tool now; defer the manager tracking panel until the governance ruling in R4 clears it; drop everything else — most of the rest is behavior, not software.
Microsoft's dashboards measure deployment. Fourier measures whether the work has changed.
LIVE APPLICATION
fourier-dg2.pages.dev
Open Fourier ↗

Fourier is a live Cloudflare app — onboard, add your own tools, and generate a 90-day plan right here.

Open full screen ↗
Built by Kylie Marcisz · Fourier & positioning, BSG engagement team
External check

Failed rollouts outran trust. Scaled ones paired AI with redesign.

Capability outran the systems built to use it confidently
Klarna
Reversed a ~700-role customer-service AI swap after quality suffered — delegation without staged trust-building fails.
IBM Watson Health
Lost hospital contracts after the tool, trained on hypothetical data, was relied on clinically — accountability was granted before visibility existed.
Amazon Hiring AI
Scrapped after it systematically downgraded female candidates — bias found only after live deployment.
ServiceNow
Its own AI maturity index fell 44→35 year over year, even as 90% of IT tickets went AI-handled — broad adoption isn't deep, governed adoption.
AI investment paired with organizational redesign
JPMorgan Chase
LLM Suite grew from 0 to ~200,000 onboarded users in 8 months via a secure sandbox model.
Moderna
750 custom GPTs built within two months, anchored by a 100-person peer-champion program.
EY
EYQ reached 400K employees and 115K monthly users, with 1,000+ field ideas distilled to 8 repeatable patterns.