The post Involve Before You Mandate appeared first on Voltage Control.
]]>
The organizations that are struggling with AI adoption have usually made the same diagnosis: employees are not using the tools enough. The response is predictable. Mandate harder. Add measurement. Tie adoption rates to performance reviews. Enforce through the management layer. Adoption rates go up. The dashboard looks better. And then, quietly, the resistance goes underground. Many employees don’t know if they’ll lose their job to AI. Few feel involved in the decisions about how AI gets deployed in their work. These two realities explain more about the failure of AI mandates than any tool-selection decision or change management curriculum. The organizations that have cracked AI adoption are not the ones that mandated better. They are the ones that involved before they mandated, and that sequence changes everything.
There is a logic to the mandate approach. Leadership has data showing that AI tools produce significant productivity gains for the people using them deeply. Leadership also has data showing that most employees are not using the tools deeply. The gap between what the tools can do and what employees are actually doing with them is money left on the table. So the reflex is to close that gap through authority. Roll it out. Train everyone. Track the dashboards. If usage is low, mandate. If mandates are not working, mandate with more specificity. What the mandate approach misses is that the resistance is not behavioral. It is psychological. The mechanism is straightforward incentive psychology. Executives who authorized the AI investment have staked their credibility and budget on the claim that AI makes the organization more productive. They need it to be true. Frontline workers who read job-loss headlines every morning need it to not be true, or at least to be less true than leadership claims. Both sides are filtering identical information through fundamentally different personal stakes. No mandate resolves that. You can compel tool use. You cannot compel the psychological safety that makes tool use productive.
Very few employees feel involved in decisions about how AI gets deployed in their work. That gap is worth holding for a moment. Most of the people who will be expected to use AI transformation tools… are not in the room when those tools are chosen. It is describing how organizations are making decisions. Eighty-eight percent of the people who will be expected to use AI transformation tools, to build their work around them, to advocate for them or resist them through daily action, are not in the room when those tools are chosen. They are handed the outcome of a decision they did not make and told to adopt it. In most organizations, this is not malicious. It is an artifact of how transformation projects are structured. A small group, usually IT and senior leadership, evaluates tools, secures a contract, builds a rollout plan, and then executes that plan on the rest of the organization. Employees are the recipients of the transformation, not the designers of it. This is not a neutral design choice. It is the choice that predicts resistance. When people feel that something is being done to them rather than with them, they comply at the minimum threshold that avoids punishment. They use the tool in ways that satisfy the dashboard while preserving their actual workflow. They wait for the initiative to lose steam, because most do. They interpret mandatory AI adoption as the latest version of a long history of top-down transformations that did not deliver what was promised to the people who had to live inside them.
The data point that most organizations misread is their own shadow adoption rate. According to Microsoft and LinkedIn’s 2024 Work Trend Index (a global survey of 31,000 knowledge workers across 31 markets), 75% of knowledge workers are already using generative AI at work, and 78% of AI users are bringing their own AI tools rather than waiting for IT-sanctioned options. They are not waiting for the rollout plan. They are solving their own problems with whatever tools they can reach. (news.microsoft.com, ‘Microsoft and LinkedIn release the 2024 Work Trend Index on the state of AI at work,’ May 8, 2024.) Organizations typically see this as a governance problem to solve. Stop the unauthorized use. Create a sanctioned list. Block the unapproved tools. The shadow-adoption numbers mean most of your workforce has already told you where the AI leverage is. The question is whether you are treating that signal as intelligence or as a compliance problem.

The word “involve” is doing a lot of work in AI transformation conversations, and most of what it is doing is insufficient. A town hall is not involvement. A feedback survey after the decision is made is not involvement. An optional AI showcase where the tools are already chosen is not involvement. Involvement that changes outcomes happens before the commitment is made. Vizient, a healthcare performance improvement organization, made this concrete. Before designing any AI-augmented roles or workflows, the company built what it calls a Persona-Based Adoption Model: a phased, empathetic approach to mapping what employees actually wanted from AI tools before rolling anything out. (Sam Fredin, Director of AI Strategy & Technology, Vizient, “Vizient’s AI Optimization Strategy Choice Cascade,” LinkedIn, Oct. 21, 2024.) The questions seem almost too simple. But they reframe the entire design problem. Instead of starting with AI capabilities and asking where they fit, Vizient started with human preferences and asked where AI could serve them. The difference in outcomes is structural. When workers feel that the transformation is built around what they said they needed, resistance looks different. They are not defending their turf from an imposition. They are trying to get the implementation right on something they asked for. This is not the same as letting employees veto every AI decision. It is the difference between involving people in problem definition and involving people in solution approval. You can involve people in the first without surrendering authority over the second, and the organizations that make that distinction get dramatically different adoption outcomes. Red Hat’s own account of its AI rollout describes the same dynamic. Engineers choose their own tools from a bounded set of options, treat AI experimentation as bottom-up rather than mandated, and the company frames the result as ‘a capability multiplier,’ not a threat. The result, in Red Hat’s own telling: an ‘explosion of internal experimentation’ and a workforce that is curious about AI possibilities rather than defensive against them.
There is a pattern that shows up consistently in organizations struggling with AI adoption. Leadership believes adoption is low because employees need more training or clearer tools. When you talk to the people doing the work, you find something different: pockets of sophisticated AI use, distributed unevenly, often invisible to the management layer above them. The practitioners using AI deeply are generating the practices that could scale. They are also, almost universally, not sharing what they know with the organization. At a Voltage Control executive dinner, Taran Lent, then leading product at Illumia, described a move that illustrates how shadow practice becomes organizational signal. He had written a five-page AI strategy paper with the model prompt pasted directly at the top, making the AI use visible rather than buried. He sent it to his parent company’s CEO with a TL;DR. The response was immediate: “This is what I mean by AI first. You guys get it.” Making the AI use legible made it safe to share. It turned a private practice into a signal the organization could see and build on. Rachel Brown, Managing Director of Innovation at CIBC Global Asset Management, made the complementary observation: a discovery sprint across every team sounds like a way to surface the real picture, but it ends up asking the people who are following the official story rather than the ones doing the work. What you actually need are domain leaders who understand the art of the possible in their specific context. Those people exist in most organizations. They are simply not in the room where the strategy is being set. The involve-before-mandate sequence requires knowing who those people are, which means you need a way to see the hallways before you design the mandate.
Mandate-then-involve creates the adoption loop most organizations are caught in: announce, train, mandate, measure, discover that measurement is gaming the metric, mandate more specifically, discover deeper resistance. Involve-then-mandate runs the sequence differently. The first move is to ask before you commit. Vizient’s persona-mapping approach is a good starting structure: find out what people want to do more of, what they’d do with more time, and what they’d rather not do at all. The answers tell you where the AI intervention has actual demand, which is where it will stick. The second move is to make shadow adoption visible rather than punishable. The three-quarters of your workforce already using AI without sanction are your best evidence of what the tools are good for in your specific context. Create a channel that makes their discoveries legible without requiring them to confess a policy violation. The policy may need to change before that channel can work. The third move is to frame the transformation as a two-way deal with explicit terms on both sides. What the organization commits to: transparency about what AI will and will not replace, involvement in decisions that change how work is designed, honest communication when the plan changes. What employees commit to: learning in real workflows, sharing what works, raising problems early when the implementation is not delivering. These three moves do not require new tools. They require a different set of decisions about who is in the room before the decisions are made.
If most of your workforce does not feel involved in AI decisions, you have not yet done the first move. The dashboard numbers are showing you adoption of the tool, not adoption of the transformation. Those are different things, and only one of them compounds. The leaders who will have functioning AI-enabled organizations in 2028 are the ones who used 2026 to ask the questions before they committed to the answers. The leaders who will be mandating harder in 2028 are the ones who committed first and are still trying to get everyone on board. Involvement is not the slow path. It is the only path that does not end in the loop. Want to explore what this looks like for your organization? Learn more about our AI transformation programs.
The post Involve Before You Mandate appeared first on Voltage Control.
]]>The post Design the Struggle Back In appeared first on Voltage Control.
]]>
Every AI rollout gets scored the same way: how much friction did it remove. Fewer approval steps. Faster first drafts. Less waiting on someone senior to review the work. The instinct is universal because it is usually right. Most organizational friction is waste, and removing it is the whole point of transformation. Some of that friction was never waste, though. It was training. The lawyer who used to draft the routine contract herself, the analyst who built the first version of the model by hand, the associate who wrote the memo nobody important would read: none of that work existed because it was the best way to produce the artifact. It existed because doing it badly, slowly, under supervision, was how a junior person became a senior one. Remove that work with AI and you have not eliminated inefficiency. You have eliminated the on-ramp. This is not a call to slow down AI adoption. Slowing down protects the wrong thing, the existing shape of the work, not the developmental value inside it. The leadership move that actually holds is different and harder. It is deliberately designing struggle back into the system AI just made frictionless. Picture a junior analyst three years into a role that, five years ago, would have had her building forecasting models from scratch for the first eighteen months. Today AI builds the first version in minutes, correctly, most of the time. She reviews it, approves it, moves on. She is faster than her predecessor ever was at the same tenure. She is also, by every account from the people who manage her, worse at knowing when the model is wrong. Nobody made a decision to let that happen. It happened because nobody made a decision at all.
When leaders notice that AI is eroding how junior people learn, the reflex is protective. Ring-fence some tasks. Tell the senior team to do things “the old way” some of the time. Add a training program to compensate for what the workflow no longer teaches. None of it works, because it treats a design problem as a scheduling problem. A training program bolted onto a workflow that AI has already hollowed out is not struggle. It is theater. The junior person knows the real version of the task is being done by AI three doors down. Practicing on a sandboxed exercise that nobody actually depends on does not build the same judgment as doing the work when it counts, because judgment is built under real stakes, not simulated politeness. There is a second failure mode that looks like caution but is really nostalgia: refusing to automate anything a junior person currently does, on the theory that all friction is developmental. That is just as costly as removing everything. Most of the friction in most workflows is genuinely waste. Formatting, reformatting, chasing down a source, restating something already said in a different template. AI should eat all of that immediately and completely. The leadership move is not protecting the old workflow and it is not automating everything in sight. It is building a new workflow that has real stakes and is still safe enough to fail inside.
The clearest version of this shift already exists in the market: the GenAI simulator, a realistic, high-stakes practice environment where the cost of a bad decision is contained but the decision itself is not simplified. We wrote about this shift in depth here, and the data is not subtle. Bank of America uses simulators to train financial advisors on the hardest conversations they will have with clients, the ones carrying real emotional and financial weight, before those advisors are in a room with real money on the line. \SOURCE: Bank of America, The Academy – [https://careers.bankofamerica.com/en-us/career-development/the-academy\] What makes a simulator different from a training module is that it preserves the actual difficulty of the judgment call while removing the actual cost of getting it wrong. A compliance officer practicing on a simulated ambiguous disclosure is still wrestling with ambiguity. A new manager rehearsing a layoff conversation with a simulated employee is still managing real discomfort, real hesitation, real second-guessing. AI can generate the finished artifact instantly. It cannot generate the discomfort of deciding under uncertainty, and that discomfort is the entire training mechanism. Skip it and you have skipped the part that actually builds the skill. Most simulators in production today are company-built, not vendor-bought, and that is not a gap to wait out. It is a signal that this capability has to be designed on purpose rather than procured off a shelf. If your organization has not identified at least one high-stakes judgment call worth building a practice environment for in the next twelve months, that is the first gap to close, and it is worth closing before the senior people who currently hold that judgment start leaving. Picking which judgment call to simulate first is not a hard problem once you frame it correctly. Ask where a bad call currently costs the most, in dollars, trust, or time, and ask who currently only learns that judgment by making the mistake live, in front of a real client or a real dataset. That intersection, high stakes and currently learned the hard way, is the shortlist. Most organizations have two or three candidates on it, not twenty, which is exactly why this is buildable in a year rather than a decade

Most automation decisions get made at the task level. Can AI write this memo. Can AI draft this contract clause. Can AI build this model. The right question sits one level up: if AI does this task, what does the person who used to do it now do instead, and does that new thing still build judgment, or did the promotion happen in title only. This is a redesign question, not a tooling question, and most organizations skip it entirely. They automate the task and assume the person “moves up the stack” without ever specifying what moving up the stack actually requires them to practice. The associate who no longer drafts the memo needs a genuinely new developmental task, something like reviewing three AI-generated drafts critically and defending, in front of someone senior, which one is right and why. Without that explicit reassignment, the associate does not move up the stack. They just stop practicing. Do this redesign work before the automation ships, not after someone notices the bench thinning out. Naming the replacement developmental task, in writing, as part of the rollout plan, is the actual substance of the redesign. Everything else, the tool selection, the pilot, the rollout timeline, is just removing a step. The step you removed has to be replaced with something that still teaches, or the organization has quietly traded a training pipeline for a productivity metric. This is also where facilitation earns its place in the AI conversation, and not as a soft add-on. Deciding which developmental tasks survive automation is a genuine disagreement waiting to happen: the people who benefit from moving fast and the people who are responsible for who exists in the role five years from now rarely agree on the trade instinctively. Someone has to structure that conversation on purpose, in the room, before the rollout, or it never happens and the default answer becomes whatever is fastest to ship. “Review and defend” is a useful default for what the replacement task looks like, but it only works if the defense is real. That means the junior person does not just initial the AI draft. They have to be able to explain, out loud, to someone who will push back, why this version and not one of the other two the model could have produced. If nobody ever pushes back, the review step degrades into the same rubber stamp the memo used to be, just with less writing involved.
The test that separates friction worth keeping from friction worth removing is simple to state and hard to apply: does this friction develop the person doing it, or does it just drain them. Formatting a document by hand develops nobody. Deciding what argument the document should make, under real constraints, with real consequences for getting it wrong, develops everyone who has to do it. AI should eat the first kind of friction completely and immediately. The second kind is the one leaders need to protect, redesign around, and in some cases deliberately reintroduce. Applying that test requires actually walking the workflow, task by task, and asking who is currently doing the developmental version of each step and whether AI just quietly took it from them without anyone deciding that on purpose. Most leaders have not done this walk. It takes an afternoon with the team that actually does the work, not a quarter with a consulting deck, and it is the single highest-leverage hour available to anyone worried about the bench five years out. The output of that afternoon should be a short, specific list: which tasks stay fully automated because the friction there was never developmental, which tasks get a redesigned developmental replacement, and which one or two tasks get deliberately protected from automation for now because nothing has been designed yet to replace what they teach. That third category should be small and it should have an expiration date. Protecting a task indefinitely is the nostalgia failure mode again, just moving slower.
None of this happens by memo. It happens in a room with the people who actually do the work, walking the real workflow step by step, naming out loud where the struggle currently lives and what it currently builds. That conversation surfaces disagreement fast: the senior person who says “that’s just busywork” and the junior person who says “that’s the only place I ever get real feedback” are describing the same task from two different vantage points, and both of them are right about their own experience. Leaders who skip that conversation and redesign the workflow from a whiteboard alone consistently guess wrong about which friction is developmental. The people doing the task know. The redesign only works if someone facilitates that disagreement into a shared answer instead of letting the loudest voice or the fastest deadline decide by default. The senior person in that room also needs to hear something uncomfortable: the busywork they are relieved to hand off might be the exact thing that made them good at their job. That is not an argument for keeping it exactly as it was. It is an argument for taking seriously what it built before deciding it is safe to remove.
Nothing about eroded apprenticeship shows up in this quarter’s numbers. Output goes up, cost goes down, and the dashboard looks great. The cost shows up later, when the senior people who have been quietly absorbing junior work retire or move on, and the organization discovers it has no one who actually built the judgment to replace them. By then the fix takes years, not an afternoon. That lag is exactly why this requires deliberate leadership action instead of waiting for the market to fix it on its own. Nobody gets punished this year for skipping the redesign. Somebody gets punished badly in year five, and the people making today’s automation decisions are rarely the ones who will answer for that later. Design the struggle back in now, while it is still cheap. The friction AI removed was never the point. The friction leaders choose to keep on purpose is.
The post Design the Struggle Back In appeared first on Voltage Control.
]]>The post Experience Starvation appeared first on Voltage Control.
]]>Here is what the productivity dashboards don’t show: every time a senior developer uses AI to write the code she would have delegated to a junior, that junior role doesn’t get eliminated in a dramatic announcement. It just stops getting backfilled. Every time a principal consultant uses AI to produce the first-pass analysis a second-year associate would have sweated through, there’s no reorganization memo. Just a slightly smaller entry class next year. This is not AI taking jobs. This is experts taking junior jobs with AI assistance. Tori Paulman, a Gartner analyst, named this in early 2026: experience starvation. “When experts use AI,” she said, “they’re able to do a lot more work. And so what happens is we see what we call experience starvation, which is that now there’s nothing easy for people to cut their teeth on.”

The mechanism is subtle. The consequences are not. And most organizations won’t see the damage until it’s too late to reverse it.
The labor market data has become hard to ignore. Erik Brynjolfsson and colleagues at Stanford’s Digital Economy Lab analyzed ADP payroll data (ADP processes payroll for more than 25 million U.S. workers, though the study’s actual working sample is 3.5 to 5 million workers per month after data-quality restrictions) and found that early-career employment in the most AI-exposed occupations declined 16 percent, controlling for firm-level shocks, since late 2022\. Developer employment among workers aged 22 to 25 is down nearly 20 percent from its peak in late 2022\.
These aren’t layoff numbers. They’re quietly empty desks. Yale School of Management’s Chief Executive Leadership Institute put it directly: “The biggest impact of Agentic AI on jobs will not be the layoffs we can see. It will be the opportunities that never materialize.” Deloitte’s 2025 Global Human Capital Trends survey, drawing on nearly 10,000 business and HR leaders across 93 countries, found that 66 percent of managers already report their recent hires are not fully prepared for their roles. That figure was in motion before AI became the dominant rationale for junior hiring cuts. Now those two trends are compounding. And PwC’s 2026 AI Jobs Barometer, which analyzed more than a billion job postings, identified a second form of the same problem. Entry-level positions in AI-exposed occupations are now seven times more likely to demand skills historically associated with experienced workers. The floor of what counts as “entry-level” has risen sharply, while traditional entry-level openings shrank 10 percent. The ladder hasn’t been removed. The rungs have. The result: recent graduate unemployment has climbed to nearly 6 percent, rising approximately twice as fast as the overall workforce since 2022, and underemployment for recent graduates sits at 42.5 percent, per the New York Fed’s Labor Market for Recent College Graduates series.
The standard case for eliminating junior roles goes like this: AI can produce better output faster than a junior employee. The economics are straightforward. This framing measures the wrong thing. Junior employees contribute to organizations primarily through their development, not their current output. The slide decks, the data cleaning, the first-draft analyses – these matter less than what they are producing in the person doing them. The work is training. The output is a byproduct. Linda Argote’s decades of organizational learning research established something most executives don’t treat as a serious operational risk: knowledge in organizations is not permanent. It is actively subject to “organizational forgetting” through employee turnover, decaying social networks, and broken pipelines. If knowledge were cumulative and stable, disrupting junior pipelines would hurt individuals but leave organizations intact. Because organizational knowledge decays when the pipeline that maintains it breaks, experience starvation is a threat to institutional capability, not just to individual career paths. This is what Paulman was pointing at with her concept of discernment: the skill of evaluating GenAI output by verifying its accuracy and judging its relevance and usefulness for the task at hand, the fourth of four essential GenAI skills she names for every worker.
Discernment is the accumulated ability to assess AI outputs. It requires having been wrong. It requires having produced something confident and incorrect and having someone with more experience show you why. It requires enough edge cases that you know when the plausible answer is the dangerous one. AI can generate plausible content at scale. Discernment determines whether to trust it. And you cannot build discernment by watching AI do work. You build it by doing work, failing, and adjusting under the supervision of someone who has already made those mistakes. When a senior engineer uses AI to produce the code a junior would have written, she produces the code. The junior doesn’t develop the discernment. The organization looks more productive in the short run and more fragile in the medium one. Ethan Mollick of Wharton made the same point in a recent New York Times roundtable on the AI workforce. “Field experience is often crucial to evaluating work you didn’t create yourself, whether it comes from humans or A.I.,” he said. A senior person can glance at a draft, a contract, or a block of code and know in seconds whether it was produced by an expert or an idiot. “If you have no experience, you can’t do those things.” The mechanism that built that experience has a name and a long track record. “We had this great technique, which was apprenticeship,” Mollick said. “It’s worked for 4,000 years.” A junior does the grunt work, a senior assesses it, both learn, everyone gets paid. “And that all collapsed” the moment AI started doing the grunt work instead.
The insidious feature of experience starvation is the lag between cause and consequence. Stop hiring entry-level talent today, and your organization does not immediately become less capable. Your senior people are still there. They are, in fact, more productive than ever. The metrics look fine. A Harvard Business School and Revelio Labs study of 62 million workers across 285,000 firms found that junior employment at AI-adopting companies declined 7.7% within six quarters of significant AI adoption, while senior employment was virtually unchanged. On the surface, the organization is stable. Beneath the surface, the pipeline has stopped flowing. When the seniors retire, or leave, or move to other organizations, the people who should have been ready to replace them aren’t there. The ones who are there have thinner experience than anyone anticipated. They have used AI to produce output. They have not been wrong in the ways that build judgment. Gartner’s own research points to where this lands. Gartner predicts that by 2027, half of companies that attributed headcount reductions to AI will rehire staff to perform similar functions, often under different job titles. Companies will discover they eliminated roles that contained more judgment-critical work than their headcount analysis suggested. Klarna is already in this cycle. The company eliminated roughly 700 customer service positions for an AI assistant its CEO said handled 75 percent of customer interactions; by early 2025, satisfaction had dropped and Klarna was rehiring for the roles it had cut.
The talent pipeline, once drained, doesn’t refill quickly. Kaelyn Lowmaster, a director analyst in Gartner’s HR practice, put it plainly: “If you’re not developing people in-house, you might have talent pipelines internally that dry up.”
Meanwhile, a Gartner survey of 110 heads of HR found that 22 percent of CHROs report at least one business leader in their organization has already stopped hiring for entry-level roles because of AI automation (Gartner, 4Q25, published July 27, 2026). Not a future risk. Something that’s already begun. Not in the future. Now.
Experience starvation is not happening in isolation. It is one of four simultaneous forces pressuring the same pipeline. Skills atrophy compounds the problem. When AI handles the foundational tasks, people stop exercising the skills those tasks built. Consider a useful image: decide to stop walking and use a scooter everywhere. Thirty days later, try to walk. The muscles have atrophied. The organizations removing junior work from their workflows are, in many cases, also removing the regular exercise that keeps senior judgment sharp. Labor scarcity adds a third pressure. The World Economic Forum’s 2025 Future of Jobs Report projects that 59 percent of the global workforce needs brand new skills within the next two to three years, with 19 percent requiring actual role changes. . This is not a future scenario. Organizations are already navigating a talent market where the supply of experienced workers is constrained. And then there is the reskilling gap itself. Companies that have cut junior pipelines for efficiency will find, in several years, that there is no internal cohort to reskill for the roles AI is creating. Those roles go unfilled or get staffed expensively from outside, with workers who carry none of the organization’s institutional context.
It would be intellectually dishonest not to name the counter-argument. Not all junior work develops judgment. NBER research found that meaningful AI employment effects are concentrated in a minority of firms: more than 90 percent of executives report no measurable AI impact on their own firm’s employment, a self-reported figure rather than a directly measured outcome. Experience starvation is a risk concentrated in organizations that are actively automating at scale, not a universal condition. More precisely: routine, codifiable work that requires no discernment can often be safely automated without talent pipeline cost. The question is not whether AI can do a task. The question is whether doing that task developed the judgment the organization needs five years from now. For a significant share of knowledge work, the answer is yes. The path from junior lawyer to senior associate goes through the research memos that got marked up. The path from analyst to vice president goes through the spreadsheets where you were wrong and got corrected by someone who’d seen it before. The path from junior engineer to staff engineer goes through the bugs you introduced and the debugging you had to do to find them. Remove those experiences and you don’t just save money. You shorten the path that would have made the next generation of senior people. The organizations that will get this wrong are the ones optimizing on output efficiency without asking what each category of junior work was producing in the people doing it.
The answer is not “don’t use AI.” It is designing deliberately for what Paulman calls developmental friction: preserving the work that builds capability even when AI could handle it. Paulman’s “Option 3” workflow is the practical starting point. Option 1 is the expert training the rookie directly, which is developmental but slow. Option 2 is the expert using AI to do the work, which is fast but eliminates the development entirely. Option 3 is the expert building the prompt or template, the junior executing the work with AI assistance, and the expert reviewing the insights and providing coaching. Nobody produces output the old way. The junior gets exposure to the decision-making layer and the feedback loop they need to build discernment. This is how Vizient approached role redesign before deploying AI into their workflows. They asked their workers: what do you want to do? What would you do with more time? What work do you hate? They built the new role design around the answers. Human-centered design applied to AI transformation produced something different from pure efficiency logic. The emerging category of GenAI simulators offers another pattern, particularly for roles where doing genuine work carries too much risk for learning purposes. Bank of America built a conversation simulator for financial advisors to practice high-stakes conversations before handling real calls. Hiscox Insurance used a GenAI simulator for certification training and found an 85 percent improvement in skills and a 75 percent reduction in certification failures. The simulator creates the difficulty, the error, the correction. It provides developmental friction without production risk. None of this is as efficient as having the senior person use AI to do everything. But efficiency is not the right metric when the thing being produced is judgment.
The organizations that get this right will have a compounding advantage that doesn’t show up in any quarterly metric. Five years from now, they will have a cohort of senior people who built their judgment through real work at lower stakes, who developed discernment through being wrong and being corrected, who carry institutional context because they were in the room when decisions were made. The organizations that optimized junior roles away to fund AI productivity will face a different reckoning. Not dramatically, and not soon. The pipeline fails quietly, and over time. You notice it when the seniors turn over and the people behind them are thinner than expected. By then, rebuilding is expensive and the institutional knowledge you assumed could be documented turns out to be harder to reconstruct than it looked from the outside. Tracey Franklin, Moderna’s Chief People and Digital Technology Officer, described converting what was “normally a junior-level HR analyst type” into a GPT. In the same month, Moderna cut 10 percent of its digital technology headcount. That equation balances on a spreadsheet. What it doesn’t calculate is what those junior analysts would have become by 2029.
Talent pipelines don’t collapse loudly. They dry up slowly, and quietly, and the cost only becomes visible when you need what they used to produce. Voltage Control helps organizations design for both the speed that AI enables and the human capability that only judgment-building friction develops. If you’re navigating this tradeoff, we’d like to talk.
The post Experience Starvation appeared first on Voltage Control.
]]>The post The Accountability Gap appeared first on Voltage Control.
]]>The data that should reframe every AI budget conversation came out this month. Gartner surveyed 350 global executives at organizations with more than a billion dollars in annual revenue, all of them running or piloting what Gartner calls autonomous business capabilities. Eighty percent have cut their workforce. Those cuts are not producing returns. Workforce reduction rates are nearly identical between organizations reporting high ROI from AI and organizations reporting low or negative ROI. Cutting is not the mechanism of value creation. It is the mechanism of budget release. Budget room is not return. And yet nobody is getting fired for missing the return. They are getting credit for the cut. That is the accountability gap.

AI transformation has a measurement problem, and it is not subtle. When an organization uses AI to automate processes and reduces headcount, there is an immediate, legible, auditable number: cost savings. That number shows up in the quarterly report. It gets attributed to the AI initiative. It earns the sponsoring executive a line in the board presentation. What does not show up is what was destroyed in the process. Three separate research efforts put numbers around the problem. MIT Project NANDA’s 2025 study found that ninety-five percent of enterprise AI pilots deliver zero measurable financial returns within six months. This number is striking
Only 28 percent of AI use cases in infrastructure and operations fully succeed and meet ROI expectations. Grant Thornton’s 2026 AI Impact Survey found that seventy-eight percent of business executives lack confidence they could pass an independent AI governance audit within 90 days. These are not statistics from organizations at the fringe of AI adoption. These are the median outcomes across organizations large enough to be writing the checks. If this were any other capital allocation category, we would call it a crisis. We would ask who was accountable. We would want to know, with specificity, what went wrong. Instead, we are giving executives credit for headcount reductions and calling it AI leadership. The problem is structural, not behavioral. Organizations have optimized their measurement systems for legibility, and layoffs are legible. You can count them, report them, attribute them to a program, and show them on a slide. The opportunity that gets destroyed in the process is not legible. You cannot count what you failed to build.
The accountability gap exists because opportunity destruction and cost savings operate on fundamentally different timelines. When you cut 30 positions, that number is immediate, auditable, and attributable. When you cut the team responsible for governing your AI systems, or reduce the people closest to the actual work who could have guided how those systems improve, the destruction does not appear on any dashboard. It appears six months later as stalled AI performance. It appears twelve months later as a system that was never adapted to the evolving needs of the business. It appears two years later as a talent base that no longer has the organizational knowledge to govern the systems that replaced them. By that point, the executive who made the cut has moved to a different role, or the organization has attributed the shortfall to new external factors, or both. The connection between the original decision and the downstream damage is no longer visible. This is not bad faith. It is a natural consequence of how organizations measure. Output accounting tracks what you produce or eliminate in a given period. It is efficient, legible, and compatible with quarterly reporting cycles. Outcome accounting tracks whether those outputs are producing the results you actually wanted. It is harder to measure, harder to attribute, and incompatible with the timeline on which most executive careers are evaluated. AI transformation is suffering from a forced adoption of output accounting applied to a problem that requires outcome accounting. You measure the cut. You do not measure whether the cut advanced the mission. And because nobody is measuring the mission, nobody is accountable for it.
Here is what makes this problem structural rather than merely behavioral: the cuts that look best under output accounting are often the ones that destroy the most value under outcome accounting. Gartner’s May 2026 human-amplified business research identifies the capabilities that determine whether an organization can sustain and expand AI performance over time. These are the people who give AI context. The people who govern how automated decisions get made. The people who adapt systems as the work evolves. The people who understand both the domain and the technology well enough to catch errors the system cannot catch for itself. These are not the roles that survive efficiency-focused headcount reduction programs. They are rarely the roles with the clearest ROI justification in a traditional cost model. Their value is mostly upstream: they prevent failures before those failures become visible, they improve systems before those systems cause problems at scale, and they build the organizational knowledge that allows technology to be used at increasing levels of sophistication over time. Consider what actually happens when an organization deploys AI to handle a function and simultaneously reduces the team that was doing that function. The AI begins operating. It does what it was trained to do. It also makes errors that the people who were just let go would have caught, because those people understood the edge cases, the organizational context, and the exceptions the system was never taught to handle. The AI does not get better on its own. It gets better when humans guide it, correct it, expand its scope, and translate domain knowledge into system improvements. Cut those humans, and you freeze the system’s capability at whatever level it was at when the cuts happened. This is opportunity destruction. It does not appear in the budget variance report. It appears in the AI initiative that was supposed to transform the business but, three years later, still does the same thing it did at launch.

The governance gap compounds this. Grant Thornton’s 2026 AI Impact Survey found that seventy-eight percent of executives cannot pass an independent AI governance audit within 90 days. This number is striking, and it is also, in some ways, the wrong metric to obsess over. Six months is not long enough for an AI initiative to transform the operating model of a complex organization. The organizations using six-month evaluation windows are setting up a measurement system that will always find AI wanting, because they are measuring a transformation initiative with an efficiency-improvement timeline. The deeper problem is that most organizations have not defined what they are building toward. They have defined what they are eliminating. The pilot documents tell you how many positions will be displaced, what the projected cost savings are, and when the payback period is expected to occur. They do not tell you what the business will be able to do in three years that it cannot do today, what human capabilities will be required to govern and expand those systems, or how the organization will develop the expertise that allows AI to operate at increasing levels of sophistication. Without that definition, outcome accounting is impossible. You cannot measure progress toward a destination you have not defined. And without outcome accounting, the accountability gap persists. Executives continue to get credit for cuts, and nobody is accountable for the opportunity quietly destroyed along the way. The governance gap compounds this. Seventy-eight percent of executives cannot pass an independent AI governance audit within 90 days. Most organizations are running AI systems without clear accountability for how decisions get made, how errors get caught, or how the system gets improved when it produces bad outcomes. The financial accountability gap is mirrored by an operational governance gap.
Gartner’s prescription runs counter to the prevailing logic of AI-driven workforce reduction. They call it human-amplified business: investing in the skills, roles, and operating models that let people guide, govern, expand, and transition autonomous systems. That is an investment argument, not a reduction argument. The organizations reporting genuine ROI from AI are not the ones that made the deepest cuts. They are the ones that built governance before they built scale. They prepared their workforce before they demanded returns. They had the discipline to stop programs that were not working, which requires having people in place who can evaluate what working actually looks like. In practice, human-amplified business looks like something specific. It looks like retaining and developing the people who understand both the domain and the data well enough to direct AI outputs. It looks like building new roles: not just people who use AI tools, but people who can evaluate system performance over time, identify drift, make judgment calls the system cannot make, and adapt processes as the technology changes. It looks like treating organizational knowledge as a strategic asset that needs active investment, not a cost to be rationalized away. The research also clarifies what happens when organizations skip this. The 95 percent failure rate on six-month ROI. The stalled governance. The talent loss. The organizations that lose their ability to govern AI systems also lose their ability to improve them. That is not a technology problem. That is an organizational design problem.
There is a practical version of this, and it starts before the first cut is made. Before reducing headcount in any AI-adjacent function, an organization should be able to answer three questions with specificity. What is this organization trying to be able to do in three years that it cannot do today? Which human capabilities are required to govern, expand, and adapt the AI systems that will support that future state? Are the people being reduced essential to those capabilities? If the answer to that third question is yes, the cut is destroying opportunity. The budget room it creates is real. The opportunity cost is also real. Both belong in the analysis, and both should be presented to whoever is approving the reduction. This is not an argument against efficiency. It is an argument for measuring efficiency correctly. Output accounting tells you what you cut. Outcome accounting tells you what you built and what you destroyed. Organizations that refuse to do both will continue optimizing for the metric that makes this quarter look good at the expense of what they are trying to become. The practical implication is a different kind of board presentation. Not “we reduced X positions and saved Y dollars through AI.” But “we reduced X positions, which freed Y dollars. Of that, we reinvested Z percent in the human capabilities required to govern and expand the systems that replaced those positions. Our outcome metrics for this initiative are A, B, and C. In twelve months, we will show you whether we hit them.” That presentation is harder to make. It is also the only one that closes the accountability gap.
Every executive running an AI transformation program is making an implicit choice between output accounting and outcome accounting. Most are not aware they are making it. Clara Shih, the former Salesforce and Meta AI executive, named the choice plainly in a recent New York Times roundtable on the AI workforce. “The key thing about A.I. agents is that they all have a goal,” she said. “And it depends on who deploys it, because whoever deploys it gets to set the goal. Maybe the goals of A.I. so far haven’t been aligned with the goals of regular people. But that’s a choice we can make.” The accountability gap is what opens up when no one names that choice as a choice. The ones who are aware ask different questions. Not “how many positions can this eliminate?” but “which human capabilities are irreplaceable in an AI-augmented operating model?” Not “what is the six-month payback period?” but “what does our AI governance look like in year three?” Not “how do we capture the cost savings?” but “how do we build the organizational competency that lets us capture value at increasing scale?” The accountability gap will close eventually. It will close when the organizations that optimized for cuts run out of runway and have to reckon with what they built versus what they destroyed. It will close when investors and boards start asking about AI governance with the same rigor they apply to AI investment. It would be more useful if it closed before either of those things happens. The question worth asking now is whether your measurement system would catch opportunity destruction before it becomes irreversible. If the answer is no, that is the governance gap worth closing first, before the next round of AI-driven workforce reductions. Voltage Control works with executive teams building the organizational structures and human capabilities required to run AI transformation at scale. If this is the conversation you are trying to have inside your organization, we can help you start it.
The post The Accountability Gap appeared first on Voltage Control.
]]>The post Lead Like a Conductor appeared first on Voltage Control.
]]>Walk into any leadership offsite and watch what the room is designed around. It is almost always an execution exercise. How do we build faster? How do we reduce cycle time? How do we ship more? That was the right question for twenty years. It is the wrong question now. AI has fundamentally changed what is worth optimizing. The execution layer, the part of your organization that turns decisions into output, is being automated at a pace that makes traditional throughput bottlenecks look like legacy concerns. Code writes itself. Reports generate in minutes. Analytical tasks that once anchored quarterly planning cycles now take an afternoon. The constraint that used to define leadership’s job is dissolving. What replaces it is not a technical problem. It is a human one. When execution collapses as the bottleneck, the new speed limit is human consensus: the time it takes for your leadership team to align on the right direction, navigate the competing priorities beneath the surface agreement, and move with enough shared conviction to act rather than stall. That is a facilitation problem. And facilitation is about to become the core leadership competency for the AI era.

For most of the history of modern management, the logic of leadership authority made sense on its face. The person who produced the best work earned the right to guide others doing it. The most technically excellent individual became the team lead. Execution quality was the primary credential. That model worked when execution was the constraint. If you were best at doing the work, you were also the most credible guide for how it should be scaled and improved. The leader’s value was embedded in their ability to produce and direct production. AI is ending that logic. A single expert, amplified by AI, can now match the output of a team. The question organizations face is no longer how to produce more. It is how to align on what to produce, and why, with the speed and fidelity that determines whether the production was worth anything at all. That requires a different skill set. Not production. Orchestration. Not executing better than everyone else in the room, but helping everyone in the room think and decide together well enough that their collective output is worth more than the sum of its parts. This is the leadership job that AI is creating. Most organizations are not yet building for it.
Joe Mariano, Senior Director Analyst / Senior Principal Analyst at Gartner, reached for a metaphor that cuts through the abstraction: digital-workplace leaders in the AI era are conductors. [Gartner, “Digital-Workplace Leaders as Conductors” (session 11e), presented at the Gartner Digital Workplace Summit 2026.] A conductor does not play every instrument. A conductor ensures proficiency across the ensemble, maintains it through rehearsal, and governs what the orchestra plays. The output is shaped not by executing the work but by designing the conditions under which the ensemble can perform at its highest level. This is a precise description of what leadership must become. The conductor’s value is not in personal output. It is in coordinating the output of everyone else toward something coherent. The conductor reads the room, feels where the ensemble is drifting, and intervenes at the level that produces the most lift. A gesture here, a structural choice before the performance begins. When something is off, the conductor does not pick up an instrument and play the missing part. The conductor adjusts the conditions until the ensemble can play it correctly. That maps directly to what leadership now requires. When execution is cheap, the leader who can produce the most output holds less competitive advantage than the leader who can align the most people around the right output, fast. The bottleneck has shifted from execution to alignment. And alignment is created through facilitation. Most organizations still have leadership development programs built on the soloist model and leadership cultures that reward individual performance. That mismatch is going to become expensive.
The word carries freight that works against it. Facilitation sounds like running meetings. It sounds like sticky notes and breakout rooms and a practitioner’s voice saying “let’s hold space for that.” That narrow version exists, and it is more valuable than most organizations acknowledge. But it is not the full picture. Facilitation, in the sense that matters for this moment, is the practice of helping groups think together, decide together, and build the shared judgment that no single person could hold alone. It is the ability to frame a decision so clearly that everyone in the room is solving the same problem rather than five parallel versions of it. It is the ability to surface the real disagreement beneath the surface-level debate, because what sounds like a tactical argument is usually a values conflict in disguise. The leader who can name that distinction is the leader who can actually resolve it. It is synthesis rather than compromise. Synthesis generates something from competing perspectives that neither perspective could have produced alone. Compromise averages them into mediocrity that satisfies no one completely. Most organizations default to compromise when they intend synthesis, and cannot tell the difference until the outcome disappoints. It is reading power and motivation in a room: who is not speaking and why, when silence signals skepticism versus deference, when to push for resolution and when to let the productive tension keep working. When to ask one more question before allowing the group to move on. It is governance: knowing which decisions are worth the room’s collective attention, which problems require human judgment and which can be delegated to the model, which questions will only get harder if avoided now. That is the conductor’s highest-value work, and it is irreducibly human. None of this is soft. It is a technical practice with learnable methods, teachable frameworks, and measurable results. And it is the practice that will determine your organization’s real velocity.
The facilitation imperative becomes specific when you recognize that AI is affecting different members of your workforce in fundamentally different ways, and each creates a distinct leadership challenge. Gartner’s workforce research maps workers across two axes: how much accumulated experience their role requires, and how much of that experience they have actually built. [Gartner, “Workforce Archetype Matrix” (session 14b), presented at the Gartner Digital Workplace Summit 2026.] Four archetypes emerge, each with different dynamics as AI accelerates. Experts hold deep domain knowledge. AI amplifies their output dramatically. The productivity gain is real and visible. The risk beneath it is concentration: Experts now absorb tasks that used to require teams, which means they also absorb the developmental opportunities that used to build the next generation. They become single points of failure wrapped in a productivity halo, and they often become them before anyone notices. Getting Experts to slow down, surface their decision heuristics, and transfer the discernment layer rather than just the procedures requires deliberate facilitation. It does not happen without a structured process designed specifically to extract what they know and make it available to others. Proteges are in complex roles but have not yet built the experience those roles require. AI creates a paradox for them. It appears to compress the path to competence, but simultaneously removes the developmental work that builds real judgment. The junior tasks they would have used to cut their teeth are absorbed by AI-augmented Experts above them. Gartner’s Tori Paulman named the mechanism directly at this year’s Digital Workplace Summit: AI is not taking entry-level jobs. Experts are. [Gartner, “AI Is Not Taking Entry-Level Jobs” (sessions 12a and 14b), presented at the Gartner Digital Workplace Summit 2026.] The facilitation challenge with Proteges is creating deliberate learning conditions in an environment that is actively optimizing those conditions away, and convincing leadership that this is worth the apparent inefficiency. Stewards are experienced practitioners whose routine work is being automated most directly. They hold institutional memory that cannot be automated, even as their current tasks increasingly can be. The facilitation challenge is transitioning them from executing routine work to governing the AI that does it, in a way that honors rather than diminishes what they have built over years. That transition is emotionally charged work. It cannot be handled with a memo. Each archetype creates a distinct consensus problem. Experts need to agree to slow down for knowledge transfer. Proteges need to be heard about what they need to learn. Stewards need real involvement in redesigning their own roles, not just notification after the decisions are made. The conductor who treats all three as the same audience will lose all three.
Taran Lent, CTO of Illumia, the higher-ed and healthcare technology company formed from the merger of Transact and CBORD, faced a version of this problem as soon as AI tools started spreading through his engineering org. Employees were experimenting individually, but that individual fluency wasn’t turning into anything the company could rely on or govern. The risk wasn’t too little AI adoption. It was adoption with no shape to it: skills and tools scattered across teams, no clear owner, no way to catch a problem before it became an incident.
Lent’s redesign started with decision rights, not tooling. He built an enablement task force explicitly designed to avoid becoming a governing bottleneck. Its job was to let people play, learn, and share what worked, rather than approve every experiment before it happened. Experimentation without any gate eventually meets reality, though, so alongside it he stood up a stakeholder review process for new AI tools and skills, a four-to-six-week approval timeline for new vendors, and guardrails built specifically to prevent incidents like an unauthenticated internal dashboard slipping into production. The task force owned the early “should we” conversation. The review process owned the “how do we roll this out safely” conversation once something was ready to scale. Two decisions, two owners, both explicit from the start.
The outcome Lent points to is not a single number. He credits a shared “humble, hungry, smart” culture, carried through this governance structure, with making the integration of Transact and CBORD into Illumia smoother than it might have been. The task force gave people room to build real fluency with AI. The review timeline and guardrails gave leadership a way to say yes quickly without finding out about a security gap after the fact. Skip the guardrails and you get the dashboard incident. Skip the permission and you rebuild the bottleneck the whole redesign was meant to remove.

Taran’s redesign points to a repeatable practice, with three components that matter most.
Decision rights need to be explicit before the conflict forces the issue. Most organizations only discover gaps in decision authority when two teams have already built conflicting work. AI accelerates this failure mode because execution is faster and misalignment surfaces sooner, often after significant effort has been spent in the wrong direction. The move is to map, in advance, who owns each category of decision, who is consulted, and what happens when the owners disagree. This is not administrative overhead. It is the infrastructure that enables fast alignment rather than repeated negotiation. Dissent protocols need to be designed in, not wished for. Most leadership cultures say they want honest disagreement and actually reward the performance of consensus. If the people in your room do not feel safe saying “I think this is wrong,” the disagreement does not disappear. It migrates to work, where correcting it is expensive. Build structures that invite dissent before decisions are finalized: pre-mortems that force articulation of what could fail, consent rounds that distinguish “I fully agree” from “I can live with this,” structured space for quieter perspectives before the dominant framing sets. These are not trust-fall exercises. They are engineering work on your decision-making process. Facilitation approach needs to match the archetype composition of the room. A session with Experts navigating a knowledge-transfer challenge needs a different design than a cross-functional session where Stewards are working through a role transition. The conductor reads who is in the room and what structure will surface the best of their collective thinking. This is diagnostic work, not template application, and it is a skill that can be learned and built deliberately. We’ve seen leadership teams cut their decision-making time by 40 to 60 percent. Not from faster tools. From fewer cycles. When groups make decisions with enough shared understanding to actually commit to them, they do not spend the following quarter revisiting the same direction. The alignment cost gets paid once, up front, through better process design. The alternative is paying it repeatedly through rework, and the bill compounds.
There is one more design choice the conductor owns, and it is the one most leaders miss. When execution collapses, it gives time back. The question almost nobody asks is where that time goes. Left undirected, it flows straight back into the existing backlog: the same roadmap, the same quarterly pressure, now executed faster. The team becomes a more efficient version of what it already was, generating more output against the same untested assumptions. Jeff Gothelf, who co-authored Lean UX, frames the failure precisely. Most organizations teach their people the AI tools and then send them back to ship more of what was already planned. They taught the team to use a hammer and expected a finished chair. The capability is real, but “capability without permission just gets absorbed by the feature factory.” What is missing is not a skill. It is permission: a protected day to experiment, a small budget that does not require three approvals, an experiment run on real data that the team is explicitly allowed to have fail. This is conductor work because it is a condition only leadership can set. Individual contributors cannot grant themselves the slack or the safety to fail. Those come from the person who governs what the orchestra plays. And the safety itself has to be redesigned for this moment. The old guardrails were built for deterministic tools that did the same thing every time. AI does not, so “safe to fail” has to be defined deliberately for work whose outputs vary, rather than assumed to carry over from the last era. The conductor who frees up execution time and routes all of it back into the backlog has not changed the orchestra’s job. They have only made it play the old score faster.
AI tools will evolve. The specific model your organization runs on today will be superseded. The facilitation capability your leaders build, the judgment about how to help groups think and decide together, is transferable across every tool change that follows. This is the argument for treating facilitation as infrastructure rather than as a support function you bring in for offsites. The organizations navigating AI transformation well share a recognizable pattern. Their leaders trust each other enough to be honest about what they do not know. They disagree productively rather than perform agreement. They move together even when not everyone is fully convinced, because they have learned how to build enough shared understanding to act without requiring unanimity. That trust does not come from a workshop. It comes from practicing the conditions that build it, repeatedly, in the actual work. The conductor builds the orchestra through rehearsal. Not by telling the musicians what to play. Mariano’s framing carries a second implication worth holding onto. The conductor also governs what the orchestra plays. In organizational terms, that is the most important leadership judgment of all: which decisions get made, which questions are worth the room’s collective attention, which problems require human judgment and which can be delegated to the model. That governance function is becoming more urgent as AI handles more of the execution work, and it is a job that cannot be automated away. When execution was expensive, leadership cleared the path. Now that execution is cheap and judgment is scarce, leadership’s job is to carry the organization’s judgment capacity forward: design the decisions that matter, surface the dissent that would otherwise stay hidden, ensure that the people who will need a skill later are getting the practice now. That is facilitation in the fullest sense. The organizations making this transition now, while execution still takes some time, are building something that will compound. They are developing the reflexes, the trust structures, and the facilitation capacity that let them move fast together when execution becomes free. The organizations that wait will still be stuck in the same alignment failures they have always had, except now the stakes are higher and the market is moving faster. Your team does not need a better AI tool. It needs a better conductor. Want to explore what this means for your organization? Voltage Control works with leadership teams to build the facilitation capability that AI transformation requires. Let’s talk about what changes when execution is no longer the bottleneck.
The post Lead Like a Conductor appeared first on Voltage Control.
]]>The post Measuring What Matters appeared first on Voltage Control.
]]>Sixteen experienced software developers sat down to work. They had access to the best AI coding tools available. They had been told, reasonably, to expect a 20 to 25 percent productivity boost. When the study ended, they believed, on average, that AI had made them 20 percent faster. They were 19 percent slower. That is from METR’s July 2025 randomized controlled trial, the most rigorous study of AI’s effect on experienced developer productivity published to date. The 39-point gap between what developers believed about their performance and what actually happened is not a rounding error. It is a measurement failure at scale. And it should be the first thing any executive reads before reviewing their organization’s AI productivity numbers.

Most organizations are measuring the wrong things. Not because their teams are careless, but because the wrong things are easy to count. Tokens consumed. Lines of code generated. Pull requests merged. Tool adoption rates. Story points completed. Time-to-first-response in customer service queues. Meeting transcripts enabled. These are the metrics that appear on AI dashboards across the enterprise right now. Every one of them produces a number. Every one of them trends over time. Every one of them can be presented in a board slide. None of them tell you whether your AI transformation is creating value. This is output accounting. It measures what happened: how much AI was used, how many tasks were touched, and how fast certain activities were completed. It does not measure whether the right things happened, whether the quality of work improved, or whether the organization can now do anything it could not do before. The appeal is understandable. Output metrics are fast, cheap, and unambiguous. In a transformation that feels uncertain and fast-moving, the comfort of a trending dashboard is real. That comfort is the problem.
The clearest proof is Klarna. In 2024, Klarna announced that AI had replaced approximately 700 customer service agents. By Q3 2025, the company claimed its AI agent was doing the work of 853 full-time employees and saving $60 million annually. The volume metrics looked excellent: faster resolution times, higher tickets-per-hour throughput, and lower cost-per-interaction. Then, quietly, in early 2026, Klarna began rehiring humans. What the output metrics had not captured was quality deterioration on complex interactions. Customer satisfaction scores on difficult cases had declined. The conversations that mattered most, the ones where customers were frustrated and needed real understanding, were getting worse. The metrics that declared success had optimized for speed in the cases where speed was least important. Klarna’s reversal is not a story about AI failing. It is a story about measurement failing. The organization tracked what was easy to track. The things that were hard to track, judgment quality on complex cases, customer trust in sensitive interactions, brand perception over time, deteriorated while the dashboard numbers climbed. This is not unique to Klarna. It is the predictable outcome of any measurement system that optimizes for outputs rather than outcomes.
Triumph Financial runs one of the largest payment networks in trucking, moving roughly $18 billion a year to carriers who often wait up to 90 days to get paid by brokers and shippers. In 2023, working with KUNGFU.AI, the company set out to speed up invoice funding with AI. It would have been easy to make speed the headline metric. Triumph didn’t.
Earlier attempts at automation, using rigid, hard-coded rules, had already failed once, rejecting too many good invoices to be useful. So this time, before anyone touted a faster approval time, the team agreed on three specific numbers to watch: disputed invoices, short pays, and write-offs. Not throughput. Not approval speed. The things that would reveal whether the model’s decisions were actually as good as a human’s, not just faster than one. Those numbers were put on a shared dashboard everyone could see, and the model wasn’t scaled up until a staged rollout, moving from historical data testing to a dark launch to a 100-day pilot, showed it held up.
Only after that did the speed numbers get to matter. And they mattered a great deal: invoice approval time fell from an average of 178 minutes to about 10 seconds, with more than $4 billion in invoices now running through the model and over half auto-approved. But the figures Triumph’s CTO Jason Heilig points to first aren’t the speed ones. Short pay is down 57 percent. Chargebacks are down 65 percent. Disputes are down 25 percent. The company didn’t get fast at the expense of getting careless. It got fast because it insisted on being careful first.
That is the inverse of Klarna. Klarna’s dashboard celebrated speed and volume while the quality of the hardest conversations quietly eroded underneath it. Triumph made quality the metric that had to clear the bar before speed was allowed to become the story at all.
Output metrics fail for a structural reason. They measure what was done. What leaders actually need to know is whether things are getting better. That distinction sounds simple. It produces completely different questions. Counting tokens consumed is easy. Asking whether the judgment underlying those tokens improved is hard. Counting PRs is easy. Asking whether the engineering organization is now capable of work it could not previously do is hard. Counting tool adoption is easy. Asking whether the team’s AI outputs are being accepted, revised, or rejected, and learning from the pattern, is hard. The organizations that stay in output-accounting mode past the early adoption phase are not being prudent. They are deferring accountability. McKinsey’s November 2025 “State of AI” report found that 88 percent of organizations now use AI, but only 6 percent qualify as high performers with measurable EBIT impact. Only 39 percent report any measurable business effect at all. The gap between having AI and benefiting from AI is the gap between output accounting and outcome accounting. High performers, in McKinsey’s data, are 2.8 times more likely to have fundamentally redesigned workflows. That is not a technology finding. It is a measurement finding: the organizations that ask deeper questions build different systems.
Output metrics answer the question “did we use AI?” The questions leaders actually need are harder and closer to the truth. The first: how many agents do you have running? This measures the scale of actual deployment, not licenses activated or employees who opened a tool. Agents running against real problems are a more honest indicator of organizational AI maturity than any adoption metric. The second: how long can those agents run without human intervention? This is a proxy for quality. An agent that requires correction every five minutes signals weak prompting, poor context engineering, or an immature tool setup. An agent that completes substantive work over hours signals a team that has genuinely learned to work with AI. Autonomy duration is a quality metric dressed as a timing metric. Anthropic’s internal research found that human interventions per Claude Code session fell from 5.4 to 3.3 between August and December 2025\. That four-month trajectory is more revealing than any adoption curve. The third: what novel work are you unlocking that was not feasible before? This is the one that matters most to leaders, investors, and boards. Not “we are shipping the same roadmap 20 percent faster” but “we stood up a customer-segmentation pipeline that had been on the backlog for two years because we could never justify the engineering cost.” Novel work is opportunity creation. It is the metric that corresponds to what organizations are actually hoping for when they invest in AI. None of these three questions are easy to put on a dashboard. That is the feature, not the bug. A metric that is easy to optimize is a metric that will be gamed. Consider what happened in the rooms where CEOs set personal token-consumption targets for their teams: staff ran purposeless prompts to hit the number. Output metrics corrupt the behavior they are meant to measure. Outcome metrics resist that corruption because they are attached to something real.

There is a second blind spot, and it sits on the other side of the ledger. The three questions above measure whether the work is getting better. They do not measure what it now costs to deliver it, and that is the other half of any honest ROI. For two decades, software ran on an economic assumption so reliable that most leaders stopped noticing it. The marginal cost of serving one more user was effectively zero. Build the product once, and the ten-thousandth user cost almost nothing more than the thousandth. Engagement was therefore an unalloyed good. More usage meant more value, more retention, more expansion, and almost no additional cost to carry it. AI breaks that assumption. Every interaction now consumes tokens, and tokens cost money in direct proportion to use. Jeff Gothelf, who co-authored Lean UX, put the consequence plainly: your most engaged users can quietly become your least profitable ones. The power user running fifty AI queries a day is the user you celebrated under the old economics and the user who erodes your margin under the new one. The dashboard that shows engagement climbing may also be showing cost climbing faster, and a usage metric will never tell you which. This is why measurement for AI cannot stop at value. It has to track cost-to-serve at the unit level: cost per successful task, gross margin per active user, model cost as a percentage of revenue. These are not finance-team afterthoughts to reconcile at quarter end. They are leading indicators of whether an AI product or workflow stays economically sustainable as it scales. An organization can be creating genuine value, clearing every outcome bar in this piece, and still be quietly building something that gets less profitable with every new power user it celebrates. The discipline is the same one this entire piece argues for. Measure the thing that is hard to see, not the thing that is easy to count. On the value side, that means outcomes over outputs. On the cost side, it means cost per successful result over raw usage volume. A serious AI scorecard holds both, because a transformation that creates value while quietly destroying margin is not a success the dashboard is equipped to catch.
The intellectual framework that comes closest to what is needed already exists. Eric Ries built it for a different context: startups trying to measure progress when traditional indicators, revenue, customers, and profitability, are all effectively zero. His answer was innovation accounting. Instead of revenue, measure validated learning. Instead of units shipped, measure hypothesis tests completed. Build, measure, learn is the loop. The goal is not to produce a big number. The goal is to reduce uncertainty faster than the competition. Nobody has operationalized innovation, accounting for enterprise AI transformation in a published, replicable form. That gap is confirmed across every major measurement research program. DORA has extended its software-delivery metrics toward AI. Anthropic has published a primitives framework that measures how AI is being used at the task level. Accenture and Wharton have built a skills-shift index tracking 150 million professional profiles. None of them answer whether the organization is learning faster, producing better work, or doing things it could not do before. The practitioner community is ahead of the published literature on this. Six consecutive executive dinners across Dallas, Houston, Boston, Boulder, Portland, and Raleigh surfaced innovation accounting independently, without anyone being prompted. In every room, leaders described the same measurement problem and reached for the same frame. In no room did anyone have an operationalized version to point at. That convergence is not a coincidence. It means the field is ready for a framework, and the gap is structural, not a matter of individual companies being slow.
The obvious counter-argument is that outcomes are unmeasurable. That output metrics are at least something, while outcome metrics are a vague ambition. This is a legitimate critique of poorly defined outcome goals. It is not a reason to abandon outcome measurement. Anthropic published the AI Fluency Index in early 2026, analyzing nearly 10,000 human-AI conversations to measure the quality of collaboration, not just the quantity. They identified 24 specific behaviors associated with effective AI use across four dimensions. Quality measurement is not an aspiration. It is an active research program producing real findings. The staged-measurement argument has something to it. Token consumption is a reasonable early proxy when the goal is normalizing AI use across a skeptical organization. Cultural adoption does need to come first. But most large organizations are past that phase now. Adoption is widespread. The question is no longer “will people use this?” It is “are we getting better because of it?” The Goodhart’s Law argument cuts both ways. Yes, any metric will eventually be gamed. That is a reason to rotate metrics deliberately, to pair quantitative measures with qualitative judgment, and to build measurement systems that are harder to optimize against than a single dashboard number. It is not a reason to accept measurements that are actively misleading. Outcomes can be defined. What novel work did your organization do this quarter that was not feasible last quarter? What percentage of AI outputs required human revision, rejection, or acceptance? How has autonomy duration changed over six months? These are measurable. They require judgment to interpret, as all meaningful metrics do. That is not a bug. That is what accountability looks like.
Auditing your current AI measurement against a single question produces clarity quickly: does each metric measure what happened or whether things got better? Tokens consumed, meeting transcripts enabled, PR counts, story points completed: these measure what happened. Replace them with counter-metrics that track quality alongside volume. If PR count goes up, track average PR complexity alongside it. If resolution time goes down, track customer satisfaction on complex interactions alongside it. At one roughly 1,000-person software company, the PR count told one story and average PR size told a different one. Both were necessary to understand what was actually happening. Then introduce the three questions as a leadership review practice. Not a dashboard, but a quarterly conversation: how many agents are running, how long can they run unattended, and what novel work did AI unlock in the last 90 days that was not on the roadmap before? The answers will be uneven and sometimes uncomfortable. That is the point. The organizations that build the discipline now, before their output metrics lock them into the wrong optimization, are the ones that will join the six percent achieving real EBIT impact. Not because outcome accounting is easy, but because it is honest.
Stanford HAI named this moment the shift from the era of AI evangelism to the era of AI evaluation. The question is not whether AI is transforming your industry. That question is answered. The question is whether your organization is learning from that transformation or simply counting it. Output accounting produces numbers that trend upward while the business-critical questions go unanswered. Klarna had great numbers until it did not. The engineering executive had a disappointing PR count until someone looked at PR size. The METR developers believed they were 20 percent faster until the data showed they were 19 percent slower. The dashboard that looks best right now may be hiding your Klarna moment. Outcome accounting is not optional for leaders who want to know what is actually happening. It is the discipline that makes the difference visible before it becomes irreversible. Want to explore what an outcome-accounting framework looks like for your organization? We work with leadership teams on exactly this question.
The post Measuring What Matters appeared first on Voltage Control.
]]>The post The Friction Discernment Test appeared first on Voltage Control.
]]>For most of the last two decades, removing friction was the whole job. Every extra click, every approval step, every handoff that made a customer wait was waste, and the work of good leadership was to find it and delete it. We built entire disciplines around it. Lean. Six Sigma. Growth. Conversion-rate optimization. Design thinking, in its most reductive form, became a hunt for anything that slowed the user down. The companies that got smoothest fastest won, and they deserved to. AI has now made friction removal nearly free. Whatever obstacle is left in a workflow, there is a tool that will dissolve it this quarter. The drafting step, the research step, the first-pass review, the scheduling back-and-forth, the synthesis of twelve documents into one. The reconciliation, the summary, the first draft of nearly anything. Gone, or going. And that is the trap. When removing friction costs almost nothing, the temptation is to remove all of it. But not all friction is waste. Some of it was the only thing developing your people. Some of it was the only thing keeping a human close enough to the work to notice when something was about to go wrong. Strip that out along with the rest, and you get an organization that runs beautifully right up until the moment it needs judgment it no longer has. The discipline leaders need now is not friction removal. That skill is commoditized; the tools do it for you. The new discipline is friction discernment: the ability to tell the difference between the friction that drains people and the friction that develops them. Remove the first kind without mercy. Design the second kind back in on purpose. Everything else in the new friction follows from getting this one distinction right.

Draining friction is friction that costs effort and returns nothing. The approval that exists because someone got burned in 2014 and no one has revisited it since. The status meeting that could have been a sentence. The reformatting of a report from one template into another. The manual data pull that a script does in a second. The form that asks for information the system already has. This friction does not build skill, protect quality, or surface insight. It just taxes attention and demoralizes the people subjected to it. AI should eat all of it, and you should let it. There is no virtue in preserving busywork, and no one should confuse what follows with nostalgia for it. Developmental friction is different. It costs effort and returns capability. The junior analyst who has to build the model by hand the first ten times, and only then earns the right to have a tool build it, because now they can tell when the tool is wrong. The reviewer is forced to articulate why an output is flawed before they are allowed to reject it. The team that has to argue its way to a decision instead of accepting the first plausible recommendation that appears on the screen. The new hire who sits in on the hard customer call instead of reading the AI summary afterward. This friction is slow. It feels like waste in the quarterly numbers. And it is the entire mechanism by which expertise, judgment, and trust get built. Here is what makes discernment hard, and why it is a discipline rather than a checklist. The two kinds of friction look identical on a process map. Both are steps that slow things down. Both show up in a time-and-motion study as cost. You cannot tell which is which by measuring duration, because duration is not the variable that matters. You can only tell by asking what the friction is producing. A leader who optimizes purely for speed has no way to see the difference, and will remove both with equal enthusiasm.
The error is not stupidity. It is a structural asymmetry in what leaders can see. Efficiency is legible. It shows up in the dashboard the week after you automate something: fewer hours, lower cost, faster cycle time, a clean line that goes the right direction. You can put it in a board deck. You can attach your name to it. Judgment loss is illegible. It shows up nowhere, for a long time. It hides inside the year-over-year improvement metrics and the reduced headcount and the deliverables that ship faster and look clean, right up until a situation arrives that needs taste, or context, or the ability to know what is not in the data. By then the people who would have caught it have either atrophied the capability or never built it at all. JoAnna Vanderhoef gave this hidden cost a name: capability debt, the widening gap between an organization’s apparent efficiency and its actual adaptive capacity. Like technical debt, it accumulates quietly and charges interest later. Unlike technical debt, most organizations are not even tracking it. They are removing developmental friction at speed, booking the efficiency, and treating the judgment that disappears as if it were free. It is not free. It is borrowed, and the loan comes due on the worst possible day.
This is not a motivational point. It is measurable. In a controlled study presented at the BIG.AI@MIT conference this year, Renee Gosline’s MIT team gave people cognitive tasks with AI assistance. In one condition, the AI made a recommendation and the person accepted or rejected it. In the other, the person first had to articulate their own reasoning, or predict what the AI’s reasoning was, before deciding. That single step took about thirty seconds. It measurably reduced over-reliance on the AI and preserved the person’s own critical thinking. Thirty seconds of deliberate friction kept the human’s judgment intact. Remove it, and the judgment quietly erodes until the day the AI is confidently wrong and no one in the room has kept the muscle to notice. The mechanism behind where this damage concentrates was formalized by a team of researchers at MIT, Yale, and Microsoft led by Mert Demirer. They studied what they call AI chains: sequences of work steps where the automatable steps are contiguous, so a human only has to verify the final output. The economic incentive is to keep extending the chain until the marginal cost of an error overwhelms the saved verification. The jobs that automate fastest are the ones where AI-suitable steps cluster together. Those are also, and this is the part that matters, the jobs where learning loops used to live. The junior who once did the research, drafted the slides, and watched a senior edit them loses three apprenticeship cycles per deliverable when the whole chain collapses into one automated unit. The work still gets done. The person stops getting made. So the friction you are tempted to remove fastest, the long contiguous chain, is frequently the exact friction that was developing your bench. Efficiency and capability erosion are not opposing forces you can balance. In the most automatable workflows, they are the same move

Before you remove a piece of friction, run it through three questions. They take a minute, and they are the discipline in practice. First: what is this friction producing? If the honest answer is nothing, it is draining friction. Remove it without hesitation. If the answer is a skill, a judgment, a relationship, or a moment where someone learns to catch what a system would miss, you are looking at developmental friction, and removal has a hidden cost you need to price. Second: who was getting developed here, and where will they get developed instead? Most automation quietly deletes an apprenticeship without anyone deciding to. If you cannot name where the replacement reps come from, you are not saving time. You are borrowing capability from your future bench, and the interest rate is high. Third: what happens on a bad day? Efficiency holds until something breaks, and then recovery runs on the slack and the judgment you preserved, not the slack and judgment you optimized away. If removing this friction means no human is left who could step in when the system is confidently wrong, the friction was load-bearing, and you are about to knock out a wall. Run a concrete case through it. A team proposes to fully automate the first draft of every client proposal. Question one: what is the drafting producing? Not just a document. It is where account managers learn the client’s business well enough to defend the recommendation in the room. Question two: if the AI drafts them all, where do new account managers build that fluency? No one has an answer. Question three: when a client pushes back hard in a meeting, who has internalized the reasoning well enough to respond? The honest run-through does not say “never automate this.” It says automate the formatting and the boilerplate, and keep new managers writing the core argument by hand until they have earned the shortcut. That is friction discernment producing a different, better decision than “remove it” or “keep it.”
Discernment becomes real when it changes what leaders actually do. Three moves follow directly. Stop automating contiguous chains to the end without asking what skill the chain was building. The most automatable workflows are exactly where capability debt compounds fastest, because they are where whole apprenticeships used to live. Automate them deliberately, and keep a human in the loop where the learning was, not only where the legal liability is. The liability checkpoint protects the company this quarter. The learning checkpoint protects it in five years. Start designing developmental friction on purpose. Route a deliberate fraction of automatable work to humans anyway, so the capability stays alive. Require a thirty-second reasoning step before anyone accepts an AI output on a decision that matters. Run novelty drills, where work that could be automated is occasionally done by hand to keep the skill warm. Sample AI outputs not for quality assurance but for drift. Bring in someone who has not been close to a pipeline to ask whether it is still doing the right thing. None of these are productivity moves. All of them are capability moves, and the point is not to make the system slower. The point is to keep it teachable. Change what you celebrate. When a team automates forty percent of someone’s job, the reflex is to bank the savings and move on. The better move, which we have watched work, is to make the freed capacity a deliberate conversation: what harder, more developmental, more human work does this person now get to do? Organizations that celebrate only efficiency teach their people that the goal is to automate themselves toward the exit. Organizations that celebrate the redeployment teach them that AI is how they grow into more valuable work. Friction discernment is not anti-efficiency. It is efficiency pointed at the right target.
When execution was expensive, leadership’s job was to clear the path: remove the blocker, approve the budget, unstick the review cycle. That job is mostly done, and the leaders still doing only that are optimizing a bottleneck that has already moved. The new job is friction discernment, and it cannot be delegated to a tool, because it is precisely the judgment about which judgments to keep. It is the one decision the AI cannot make for you, because the AI’s entire bias is toward removing friction, and the question in front of you is when not to. The organizations that get this right will look slower for a few quarters and less impressive in the efficiency reports. They will also still have, when the situation changes, the people who can do the work the model cannot. The organizations that remove every obstacle they can afford will discover, on the worst possible day, that they removed the ones holding the building up. Stop removing every obstacle. Learn to tell the difference. Remove the friction that drains your people. Design the friction that develops them. That is the discipline, and everything else in the new friction follows from it.
The post The Friction Discernment Test appeared first on Voltage Control.
]]>The post Trustworthiness Is Not Trust appeared first on Voltage Control.
]]>Your organization has done the work. You have accuracy benchmarks, SLA guarantees, pilot results, case studies with named clients and documented ROI. Your vendor has third-party audits. Your legal team reviewed the data handling. Your IT team certified the security posture. And your employees still are not using it. This is not a failure of evidence. It is a category error. You have been building trustworthiness. You needed to be building trust. These are not the same thing, and conflating them is why most enterprise AI adoption efforts stall at exactly the moment they should be accelerating.

Trustworthiness is what the evidence shows: accuracy rates, compliance certifications, SLAs, pilot results, and audit trails. It is an attribute of the AI system itself. You can measure it, document it, and present it in a deck. Trust is different. Trust is a psychological act that happens inside a person. It is the moment someone decides to rely on something they cannot fully verify. And that decision is not primarily driven by evidence. It is driven by experience, context, identity, and social proof from people they respect. The distinction matters because the interventions are completely different. Loading more evidence into your adoption campaign, another case study, another ROI breakdown, another compliance badge, does not move the needle on trust if the underlying psychological conditions are not met. You are solving for trustworthiness while employees are asking a different question. The question is not “Is this AI trustworthy?” The question is “Do I trust this AI, here, in my role, for this kind of work?”
Here is the pattern that reveals this most clearly. A knowledge worker who hails a robotaxi and lets software navigate them through city traffic at 40 miles per hour is the same person who refuses to accept a Copilot-generated first draft without rewriting it from scratch. Objectively, the stakes do not compare. A robotaxi error could injure them. A hallucinated summary wastes fifteen minutes. But their trust behavior inverts what the evidence would predict. Why? Because the psychological conditions are entirely different. With the robotaxi, the role boundaries are clear. The car drives; they sit. The system is visibly working in real time. Social proof from colleagues who have used it accumulates passively. And critically, their professional credibility is not on the line. If the robotaxi takes an odd turn, they observe it. They do not own it. With Copilot, everything changes. The output lands in their document, under their name, in their domain of expertise. If the summary is wrong and they forward it, that is their error. The AI did not fail. They failed to catch the AI failing. Their reputation as someone who knows their material is at stake in a way it simply is not when they are a passenger. Trustworthiness is similar across both systems, or arguably higher for Copilot given its output transparency and audit trail. Trust diverges completely because the psychological stakes differ. This is not irrational. This is exactly how trust works. The lesson for AI leaders is specific: the trust gap your employees have with enterprise AI is not primarily about the model. It is about the context in which they use it and what failure costs them professionally. An employee who trusts AI to help draft internal updates may not trust the same AI to help draft client recommendations, even if the capability is identical. The context changes the psychological stakes. The psychological stakes change the trust response. Treating both contexts as equivalent, and responding to the skepticism in the second context with more evidence from the first, is the mistake most adoption programs make.
The conventional response to adoption resistance is to produce more evidence of trustworthiness. Refine the accuracy stats. Commission an independent audit. Write up a case study from a similar organization. Schedule a lunch-and-learn to walk through how the model works. This is understandable. It is also almost always wrong. Craig Roth at Gartner’s Digital Workplace Summit named what actually happens: organizations deploy AI rapidly, loading employees with technical information about the system, and create trust deficits precisely because speed and data-loading leave no room for the gradual, experience-based trust-building that works. Speed is a trust deficit. Evidence is not trust. Research on what actually drives psychological trust identifies three factors: perceived ability (can the system do what it claims?), benevolence (does it act in the user’s interest?), and integrity (does it behave consistently and honestly?). Evidence addresses ability. It barely touches benevolence and integrity, which are primarily established through direct experience, not documentation. Worse, detailed technical explanations often activate a risk mindset rather than a trust mindset. Walking through the training data surfaces concerns about bias. Explaining confidence intervals surfaces concerns about accuracy in edge cases. Describing the audit methodology surfaces questions about what the audit did not cover. You have made the system more transparent, which improves trustworthiness. You have also made the failure modes more vivid, which suppresses trust. The information is accurate. The effect is the opposite of what you intended. This is the core tension: the moves that build trustworthiness and the moves that build trust operate through different mechanisms. Most organizations invest heavily in the former and wonder why it does not produce the latter.

Trust in AI builds the same way trust in anything builds: through repeated exposure, positive experience, social modeling, and calibrated stakes.
Small starts, visible wins.
The organizations seeing genuine AI traction are not the ones who launched enterprise-wide mandates backed by polished training programs. They ran tight pilots in one team, let people experiment with low-stakes tasks, and let word of mouth carry the initial momentum. When someone uses AI to draft a rough first cut of a weekly update and it saves them an hour, they tell people. That conversation transfers more trust than any case study. You cannot engineer the conversation directly, but you can create the conditions for it: start small, start where the AI clearly succeeds, and give people room to discover it themselves.
Top-down permission, bottom-up testimonials.
Both matter, and they serve different functions. Leadership commitment, when an executive uses AI visibly in their own work and says so, creates permission. It signals that experimentation is safe and that the organization values the output even when it is imperfect. Bottom-up testimonials from actual practitioners, not trainers or IT leads but respected domain experts who talk about specific ways AI helped them, create desire. They answer the question employees are actually asking: “Does this work for someone like me?” Top-down without bottom-up is a mandate. Bottom-up without top-down is shadow AI, happening outside governance, invisible to the organization and to any accumulated trust benefit. You need both.
Sequence use cases to build a track record.
Not all AI use cases carry the same trust-building or trust-destroying potential. A rough first draft on an internal update is low stakes and often succeeds visibly. A client-facing analysis output is high stakes and will be scrutinized in ways that compound skepticism if it fails. The sequencing of early experiences matters enormously. Start where the AI succeeds clearly, with work that is iterative and internal, where recovery from error is easy and the user stays in control. The trust you build in those contexts transfers to harder ones. The distrust from an early public failure also transfers, and it spreads faster.
Address the identity question directly.
This is the piece most adoption strategies miss entirely. Knowledge work AI almost always has professional identity stakes embedded in it. Am I still the expert if the AI writes the first draft? Am I still the analyst if the AI runs the summary? Am I still the strategist if the AI builds the framework? These questions are not irrational objections to be overcome. They are prior to the trustworthiness question. Answering “yes, the model is 94% accurate” does not address “yes, you are still the expert.” The leaders building AI fluency that holds are addressing this directly. Not by reassuring people that their jobs are safe, which most people do not believe, but by making the new shape of expertise visible. Directing a tool well is a skill. Editing a first draft to a high standard is a skill. Knowing when to override the output is a skill. Recognizing what the AI missed requires knowing your domain deeply. The expert who uses AI well is more capable, not less. Making this visible and valued is not an HR exercise. It is a trust-building move.
If you have an adoption problem, the instinct is to add more proof. Resist it. Ask instead what is driving the psychological conditions that make trust difficult. The answers are usually specific. If employees feel AI is happening to them rather than with them, involving them in defining what good output looks like makes a material difference. People trust systems they helped shape more than systems deployed at them. Letting practitioners set the quality bar for what AI-assisted work needs to meet is not a political gesture. It changes how they relate to the system. They will defend the standard they set. If the social modeling in your organization is skeptical, one champion in the right position is worth more than a hundred case studies. Not a trainer. Not an IT lead. A respected domain expert who uses AI openly and talks about what it changed for them. Their credibility transfers to the tool. If you are building the business case around accuracy stats and ROI figures, understand that you are building a trustworthiness case. That case needs to be made to procurement, to legal, to the board. It is not the case employees need. The case employees need is not about whether the AI is reliable. It is about whether they can rely on it, in their context, for their work, in a way that protects their credibility rather than threatening it.
Trustworthiness is table stakes. It gets you through the governance gate and into the pilot. Trust is what gets you to adoption. Adoption is where the value is. The organizations that figure this out are not producing better case studies. They are building different conditions: room to experiment without professional exposure, social proof from real practitioners, early use cases where success is obvious, and an explicit reframing of what expertise looks like when AI is in the room. The ones that do not will keep wondering why employees who nodded through the AI launch presentation still open a blank document and start typing. If you are working through this in your organization and want to talk about where you are stuck, the gap between trustworthiness and trust is usually where the most interesting questions live. Start that conversation here.
The post Trustworthiness Is Not Trust appeared first on Voltage Control.
]]>The post AI Is Not a Legal Shield appeared first on Voltage Control.
]]>Joe Mariano said something at the Gartner Digital Workplace Summit and it should be on a poster in every AI governance committee in the country. “AI is a tool. It is not a legal shield.” Two recent cases prove him right. In both, an organization deployed an AI system, the system did exactly what it was designed to do, the organization got sued, and the organization lost. Not because the model misbehaved. Because nobody on the human side was watching. These are not abstract risks. They are the first two real precedents we have for AI liability in the enterprise, and they both turned on the same thing: a monitoring and data gap that the legal system treated as the company’s responsibility, not the AI’s. If you are a leader thinking about AI governance, you are no longer thinking about it in theory. You are thinking about it inside the ruling that other companies have already lost.

In November 2022, Jake Moffatt visited Air Canada’s website to book a last-minute flight to attend his grandmother’s funeral. He asked the airline’s AI chatbot whether bereavement fares were available. The chatbot told him yes, that he could book the flight at full fare and apply for a bereavement refund within ninety days of travel. He booked. He flew. He filed for the refund. Air Canada denied the claim. The actual policy required bereavement-fare requests before travel, not after. The chatbot had hallucinated a refund window that did not exist. Moffatt took the airline to the British Columbia Civil Resolution Tribunal. Air Canada’s argument was the part that should make every AI governance lead pay attention. The airline argued that the chatbot was, in their words, “a separate legal entity that is responsible for its own actions.” The tribunal rejected that argument flatly. In Moffatt v. Air Canada, 2024 BCCRT 149, the tribunal ruled that Air Canada was responsible for all information on its website, regardless of whether it came from a static page or a chatbot, and ordered the airline to pay damages. The decision is short, the legal reasoning is clean, and the precedent is simple: if your AI tells a customer something false, your company said it. The chatbot does not have its own lawyer. It does not have its own bank account. It does not have legal standing. It is a tool you deployed, and the output is your output. Air Canada’s failure was not that the chatbot hallucinated. Hallucination is a known property of generative systems, and any organization deploying one in a customer-facing context should plan for it. The failure was that nobody checked. There was no monitoring layer, no review pipeline, no human in the loop verifying that high-stakes policy claims matched the airline’s actual policy. The model behaved as models behave. The organization behaved as if the model would not.
In August 2023, the U.S. Equal Employment Opportunity Commission settled a case against iTutorGroup, a tutoring company that had used an AI-driven recruitment system to screen applicants for tutor positions. The system was configured to automatically reject women aged 55 and over and men aged 60 and over. More than two hundred qualified applicants were filtered out before any human ever saw their applications. The EEOC argued, and iTutorGroup agreed in a consent decree, that the company had violated the Age Discrimination in Employment Act. iTutorGroup paid $365,000 in damages and committed to anti-discrimination training and oversight changes. It is widely cited as the first EEOC enforcement action targeting algorithmic discrimination, and it set the regulatory tone for what was to come. The interesting thing about this case is that the AI did not malfunction. It did exactly what its rules told it to do. Somewhere in the configuration, somebody had set age thresholds. The AI applied them. Hundreds of times. What was missing was the question of whether anybody should have set those thresholds in the first place. There was no review of the screening logic against employment law. There was no monitoring of who was being filtered out and why. The data quality, the rule design, the oversight layer, all of it sat inside an AI deployment that nobody thought needed governance because the AI itself was working. That is the iTutorGroup pattern, and it is more dangerous than the Air Canada pattern because it does not look like an AI failure. It looks like an AI success.
Joe Mariano walked through both of these cases at Gartner DWS 2026, and the framing he landed on is worth repeating: the failures here were not in the technology layer. They were in the layer above the technology, where humans decide what the AI is allowed to do, what data it sees, and who is watching when it is doing it. The Air Canada chatbot worked as a generative chatbot works. It produced a plausible answer to a question it did not have grounded knowledge to answer. The failure was that the airline deployed it on a high-stakes policy page without a verification pipeline. The iTutorGroup recruiter worked as a rules-based filter works. It applied the configuration it was given. The failure was that the configuration had been set by humans without legal review, and there was no monitoring on the output to flag the discriminatory pattern. Both failures, in other words, traced to the same place: the human-decision system around the AI was not designed. The technology was deployed faster than the governance scaffolding around it could catch up, and the legal exposure that resulted was real. This is the part most AI governance conversations skip. They focus on the technology, on which model, on which vendor, on which compliance certifications, when the actual exposure lives in the workflow. Who reviews high-stakes outputs before they go to customers. Who audits the rules the AI is using. Who is watching for patterns that look fine inside the model but look discriminatory in aggregate. The Gartner data backs this up. There are over 1,000 proposed AI rules and regulations worldwide right now, and not one of them has the same definition of AI. The regulatory landscape is going to get harder, not easier. Companies that are still treating AI governance as a policy document, rather than as an active facilitation problem inside their organization, are going to keep producing the next Air Canada and the next iTutorGroup.
There is a comforting fiction that some leaders are still telling themselves about AI deployment, which is that the model carries some of the liability. It does not. Across multiple jurisdictions, in multiple legal frameworks, the rulings are converging on the same answer: the deploying organization is accountable for the output, full stop. This makes sense the moment you say it out loud. The AI did not sign a contract with the customer. The AI did not file a lawsuit. The AI did not get sued. Your company did all three of those things. The model is a tool the company chose to deploy, and the output of that tool is the company’s output, the same way that an internal email written by a junior employee is the company’s email. What this means in practice is that the conversation about AI governance has to move from “is the model trustworthy” to “is our deployment of the model accountable.” Those are different questions. A trustworthy model deployed without governance is still a liability. An imperfect model deployed with rigorous governance is, in many cases, fine. The trustworthy-but-ungoverned configuration is what produced both cases above. The Air Canada chatbot was, by industry standards, a perfectly normal AI product. The iTutorGroup recruiter was, by configuration standards, perfectly capable of being used legally. Neither model was the problem. The deployment around it was.

If the technology is not the gap, what closes the gap? Mariano’s session offered four patterns, and they map cleanly to what we see when we walk into client governance work. Brain-first deployment. Before AI is brought into a workflow, the team uses human judgment to define the goal, the boundary, and the success criteria. The AI is then brought in to assist a human-defined process, not to replace the human-definition step. Air Canada skipped this. The chatbot was deployed on a policy page without anyone defining what counted as an acceptable policy answer. Human in the loop for quality control. Some volume of AI output gets reviewed by a human before it goes to a customer or a decision. The exact percentage depends on the stakes, but the principle is non-negotiable: zero human review on high-stakes outputs is a deployment, not a governance posture. iTutorGroup ran a recruitment AI with apparently no auditing of the rejection pattern. That is the failure case. Data quality management. AI systems accessing wrong, stale, or biased data will produce wrong, stale, or biased outputs with full confidence. Both Air Canada and iTutorGroup had data quality problems at the root. The chatbot was answering questions about a policy it had not been grounded in. The recruiter was applying rules that had not been audited against current law. Neither case was a model problem. Both were data problems wearing model clothing. Continuous skill and process maintenance. Governance is not a one-time training. It is an ongoing practice, with periodic reviews, audits, and skill refreshes for the people running the system. The model evolves. The regulations evolve. The use cases evolve. A governance framework that was designed twelve months ago and has not been touched since is, by definition, stale. These four patterns are not novel. They are the basic discipline of any high-stakes deployment, applied to AI. What is new is that the legal system is now treating them as the standard of care, and organizations that ignore them are losing the cases.
Here is the move most organizations miss. AI governance is not a document. It is a set of ongoing agreements between security, legal, business, and operations about what the AI can do, who is watching, and what happens when something goes wrong. Those agreements have to be negotiated. They cannot be written by one team and handed to the rest. This is where the work gets uncomfortable, because it requires the same cross-functional conversation that most organizations are structurally bad at. Legal does not want to talk to Engineering. Security does not want to talk to Marketing. The business unit that wants to deploy the chatbot does not want to slow down for a review. And so the governance conversation never happens, and the deployment goes out, and somebody loses a tribunal. The organizations getting this right are the ones that treat AI governance as a facilitated, recurring practice, not a sign-off process. They have a standing forum, with the right people, that meets often enough to keep up with what is being deployed. They produce decisions, not policy documents. They review the deployments that have shipped. They ask, every time, what would happen if this output ended up in front of a regulator or a tribunal. That is the New Friction. AI eliminated the old friction, which was execution time. The new friction is the human-decision layer that has to keep up with what AI now lets you ship. Organizations that do not invest in that layer ship faster, get sued more, and lose the cases. Organizations that do invest in it ship slightly slower, ship better, and stay out of the tribunal. If you are building or refreshing your AI governance posture right now, the question is not which model you trust. The question is which decisions your organization can keep up with, and which conversations you are willing to have to keep up with them. That is the work. That is the entire work. If your organization is in the middle of that conversation, or trying to start one, that is where Voltage Control comes in. Read our New Friction primer for the full framework, or reach out if you want to talk about where your governance posture is stuck.
Can companies be held liable for AI mistakes?
Yes. Both the Air Canada and iTutorGroup cases establish that the deploying organization is responsible for AI output, regardless of whether the output came from a human or an AI system. Air Canada explicitly argued that the chatbot was a separate legal entity. The tribunal rejected the argument. Across jurisdictions, the legal direction is consistent: the company that deploys the AI owns the consequences.
What happened in the Air Canada chatbot case?
A customer asked Air Canada’s chatbot about bereavement fares. The chatbot hallucinated a refund policy that did not exist. The customer relied on it, booked the flight, and was denied the refund. He took the airline to the British Columbia Civil Resolution Tribunal, which ruled in Moffatt v. Air Canada, 2024 BCCRT 149, that the airline was responsible for the chatbot’s output. Air Canada paid damages.
How do organizations govern AI systems effectively?
Effective AI governance is a recurring facilitation practice, not a static policy document. The organizations doing this well bring legal, security, business, and operations into a standing forum that meets often enough to keep up with deployments, audits real outputs, and produces decisions. The four operating patterns are brain-first deployment, human in the loop for quality control, data quality management, and continuous skill maintenance.
What is AI accountability in the workplace?
AI accountability is the principle that the organization deploying the AI is responsible for what the AI produces. That responsibility cannot be delegated to the model, the vendor, or the AI system itself. It lives with the humans who decided to deploy the system, configured it, fed it data, and chose how much oversight to give it.
Who is responsible when AI makes wrong decisions?
The deploying organization. In every major AI liability case to date, including Air Canada and iTutorGroup, the courts and tribunals have held the company responsible for the AI’s output. The model is treated as a tool, and tool failures attach to the operator, not to the tool.
The post AI Is Not a Legal Shield appeared first on Voltage Control.
]]>The post The Five-Year Gap appeared first on Voltage Control.
]]>“We’re worried because there are fewer entry-level jobs right now, and in five years, there will be fewer intermediate or senior-level designers. There’s going to be a gap.” That is a working landscape architect, one of 722 interviewed for a new Cornell study presented at the BIG.AI@MIT conference last month. The quote is not about a dystopian future. It is about what the practitioner is watching happen, month by month, in her own firm. The headline narrative on AI and work is about entry-level job loss. That is real, and it matters. But the more consequential story, the one that is almost invisible in quarterly earnings calls, is quieter and slower: the people who would have been senior in five years are not getting trained now. The junior did not lose the job. The junior lost the reps.

Jose Antonio Guridi and Cristobal Cheyre, researchers at Cornell, spent the last eighteen months studying how landscape architecture firms across North America are adopting generative AI. They did 25 semi-structured interviews, spent time observing operations at a prominent firm, and ran a survey of 722 practitioners. That is one of the largest datasets on AI adoption in a real profession that has been published. Three findings stand out. First, the adoption is uneven in a specific pattern. Juniors are driving it. Seniors are holding the judgment that decides whether the AI output is right. In firms that have not designed for this reversal, the junior uses AI to produce something the senior reviews. The senior edits the output. The original work the junior would have done, the intermediate steps where skill used to form, quietly disappears. Second, most of the adoption is hidden. 73% of the practitioners who use AI at work do not disclose that use to their firm. They use personal devices. They treat restrictive firm policies as an obstacle to route around rather than a signal to stop. The work gets done. The firm does not know how. Third, the firms that are handling this well have all done the same thing: they have made the adoption explicit. Structured workshops. Shared documents. Senior oversight built into the workflow, not as policing, but as a design feature. The distinction between the three patterns is where the story is.
Guridi and Cheyre name three adoption patterns. Each one produces a different organization five years from now. Passive adoption. AI arrives through software updates. The design tool adds a new button. The email client starts suggesting full paragraphs. The research database surfaces AI-generated summaries above the actual sources. Nobody decided. The practitioners absorb the change as background noise. Skill formation is whatever it would have been, minus the steps the software now does automatically. Passive adoption is the modal case. Most organizations are in it right now and do not realize it. Hidden adoption. The firm has a restrictive AI policy. The practitioners need to produce the work anyway. They open ChatGPT on their phones, paste the brief, and keep the output in their personal notes. They know they are not supposed to. They do it because the alternative is not doing the job. The 73% disclosure-gap statistic is this pattern, captured at scale. Hidden adoption looks like conformity from the outside. From the inside, it is an underground apprenticeship running in parallel with the firm’s official one, except the underground one is entirely unsupervised and invisible to every senior person who might intervene. Explicit adoption. The firm has decided, out loud, how AI fits into the work. There are designated workflows where AI is expected. There are designated workflows where AI is not welcome. There are senior reviews built into the AI-assisted paths, not as gates, but as teaching moments. Juniors get exposure to the AI-generated output. They also get exposure to the senior’s reasoning about why the output is right or wrong. This is the only one of the three patterns that preserves apprenticeship.
A team of economists at MIT, Yale, and Microsoft, led by Mert Demirer, gave this phenomenon a structural name. They call it AI chains.
An AI chain is a sequence of production steps in which each automated step flows into the next without a human in the middle. Verification happens once, at the end of the chain. The economics are obvious: verification is expensive, so fewer verifications are better. Organizations will push toward longer chains whenever the AI is good enough.
The consequence is that jobs where AI-suitable steps sit next to each other are the jobs where chains form fastest. Research, drafting, and rendering are adjacent. So are summarization, synthesis, and first-pass review. Chain the three together and you have converted what used to be a six-hour junior assignment into a thirty-second prompt and a five-minute senior review. The efficiency gain is real. So is the apprenticeship cost, which does not show up anywhere on the quarterly report.
In landscape architecture, Guridi and Cheyre watched this happen inside the firm they observed. Rendering production used to be the junior’s job. It was slow, iterative, and humbling. You started something, showed it to a senior, received criticism, started again. After two years, you had internalized the senior’s taste. After five years, you had your own.
The rendering step is now in a chain. The junior writes a prompt. The AI produces four variations. The senior picks one and edits it. The junior has watched, but has not done. The internalization does not happen the same way. The taste does not form. One practitioner put it to the researchers this way: “If you’re using something to generate everything, you miss all of these moments to be iterative and review your own work.”
The pattern is not unique to design. In May 2025, Moderna’s Chief People and Digital Technology Officer Tracey Franklin described to the Wall Street Journal a system of more than 3,000 internal GPTs, including a broad HR GPT that routes employee questions to specialized GPTs for performance management, equity, and benefits. Her own description of the workflow: “It’s like your virtual HR, AI agent. It’s what would normally be a junior-level HR analyst type, we’ve now converted into a GPT.” Same chain. Different industry. The intermediate work that an HR analyst would have done on the way to becoming a senior HR partner is gone.
The reason this pattern is so hard to see at the executive level is structural. It is not a failure of leadership attention. It is a failure of legibility. The metrics you have are the metrics that matter. Revenue per employee. Project cycle time. Client satisfaction. None of these show apprenticeship. All of them might actually improve in the short term when AI chains form, because the outputs ship faster and the staff count drops. The disclosure gap compounds the invisibility. 73% of AI users are hiding the use from the firm. Senior leaders cannot see what they cannot see. The firm’s governance layer is responding to a world where AI use is still occasional. The actual daily reality has moved past that. And the time horizon is precisely the wrong length. Five years is long enough that the consequence is somebody else’s problem, probably the problem of whoever succeeds today’s CEO. Five years is short enough that the seniors who exist today will still exist and can still cover the gap, right up until they retire. We name this pattern “Experience Starvation,” after the term coined by Gartner’s Tori Paulman at last year’s Digital Workplace Summit. Experience starvation is what you get when the workflow around the AI strips out the intermediate work the junior used to do on the way to becoming the senior. The organization continues to function. The talent pipeline quietly thins. Paulman’s framing has a sharp corollary: AI is not taking entry-level jobs. Senior people are.

The explicit-adoption firms in the Guridi study are not slower. They are not abstaining from AI. They have just designed the adoption so that apprenticeship survives. The most teachable pattern in the research is one Paulman calls the Option 3 workflow. It has three moves. The expert builds the template. The senior practitioner, who has the taste, captures her reasoning in a reusable form. The template is the artifact. It encodes the judgment. The rookie executes with AI. The junior runs the template, feeds it the project context, and gets the output. They see the template working. They see where it breaks. They do the adaptation work the template did not cover. The expert reviews the insights. The senior does not edit the output. The senior reviews the judgment the junior exercised when the template was not sufficient. The feedback is on reasoning, not on rendering. The workflow preserves three things at once. The firm gets AI leverage on the routine work. The junior gets exposure to the senior’s reasoning, not just the senior’s output. The senior spends her scarce time on the decisions that only she can make. This is what Guridi and Cheyre observed in the firms that were explicit about their AI adoption. It is not a program. It is a set of working conventions that the senior partners enforced because they had decided, out loud, that training the next generation was part of the firm’s product. The firms that had not made that decision were not using any of this. They were using AI chains that removed the work and the learning together.
Three moves that do not require a transformation program. Make disclosure safe. The 73% who hide AI use are not malicious. They are responding to incentives you set. If the penalty for disclosing AI use is higher than the penalty for hiding it, you will get hiding. Change the incentive. A one-line policy update (“we encourage AI use in designated workflows; here is how to propose a new one”) can move the whole distribution. You cannot design around a pattern you cannot see. Route some work through juniors even when AI could do it. Not all of it. Some. The criterion is whether the work teaches something the junior needs to know in five years. If the answer is yes, the junior does it. The efficiency loss is the training budget, reclassified. You are already paying for training; now you are spending it on practice instead of on certificates. Audit your senior bench replacement rate. Not headcount. Replacement rate. For every senior who will retire or exit in the next five years, who is on track to replace them? If the answer is “unclear,” you have the gap already. The only question is whether you find out now, when you can still do something, or in three years, when your best seniors are announcing and the bench is empty. None of these require new hires. None require new tools. They require the decision to design for apprenticeship at a moment when every incentive is telling you to optimize it away.
The five-year gap is not a forecast. It is a trajectory measurement. The apprenticeship loss is happening now. The consequence is scheduled to arrive in 2030\. The organizations that will have the senior bench they need in 2030 are the ones that decided, in 2026, that apprenticeship was a design problem. They built Option 3 workflows. They made disclosure safe. They kept routing work through the junior even when the AI was right there and faster. The organizations that will have the gap in 2030 are not doing anything wrong, exactly. They are optimizing for the metrics they have. The metrics they have do not measure apprenticeship. Apprenticeship erodes silently. By the time it shows up as a capability gap, the people who could have been trained have moved on to firms that trained them. The juniors are not losing their jobs. They are losing the work that would have made them senior. That is a different problem, and it hides better, and it bills later.
If your organization is watching AI chains form and is not sure whether apprenticeship is surviving, three places to go deeper. Talk to us. We help leadership teams design the workflows that keep AI leverage without losing the learning cycles. Learn more Our pillar page lays out why apprenticeship loss is one of the new frictions AI has relocated into the center of your organization. Build the capability. Our facilitation certification teaches the skills senior leaders need to run Option 3 workflows at scale.
Is AI taking entry-level jobs?
The headline narrative says yes, but the more consequential pattern is different. AI is enabling senior workers to do entry-level work themselves, which removes the on-ramps for skill development. The junior role often still exists; the work that used to fill it has been compressed into AI chains. The Cornell study of 722 practitioners shows this pattern clearly. The junior did not lose the job. The junior lost the reps.
What is experience starvation in AI adoption?
Experience starvation is a term coined by Gartner’s Tori Paulman to describe the systematic removal of the low-stakes, high-repetition work that builds professional judgment. When AI handles the steps where skill used to form, the junior misses the iterative cycles that produce taste. The organization keeps shipping. The talent pipeline quietly thins. By the time the gap shows up, the people who could have been trained are five years past the moment when training mattered.
How does AI break the apprenticeship model?
AI chains the production steps where junior workers used to learn. Research, drafting, rendering, review: each one used to be a discrete moment where a junior practiced and a senior critiqued. When those steps chain together, the junior writes a prompt and the senior edits the final output. The intermediate work, where taste forms, disappears. Most organizations have not noticed because their metrics do not measure apprenticeship. Cycle time and revenue per employee actually improve in the short term.
What is the Option 3 workflow for AI in the workplace?
The Option 3 workflow, also from Paulman’s research, has three moves: the expert builds a reusable template that encodes her judgment, the rookie executes the template with AI on real project context, and the expert reviews not the rendered output but the reasoning the rookie applied when the template was insufficient. It preserves AI leverage on routine work while giving juniors exposure to senior reasoning. It is the only workflow pattern in the Cornell research that survives apprenticeship.
How do you protect your talent pipeline from AI-driven erosion?
Three moves: make AI disclosure safe so you can see what is actually happening (the Cornell data shows 73% of users hide their AI use from their employers); route some work through juniors even when AI could do it, with the criterion being whether the work teaches something the junior needs in five years; and audit your senior bench replacement rate, not headcount but replacement rate, so you know where the pipeline is actually broken before it shows up as a capability gap.
The post The Five-Year Gap appeared first on Voltage Control.
]]>