If you use AI at work, you’re being watched.
When you open a chat window or coding harness on your company account and type in some rough thoughts, “dumb” questions, or even some personal information to get to work on a fuzzy task, go back and forth with AI over a number of turns, modify the work product until it’s done, all of that data is stored and subject to search - and therefore analysis - by multiple parties.
You probably knew or suspected this, but to lay it bare: Microsoft says Copilot “prompts and responses are captured in the unified audit log,” stored in your mailbox, searchable in eDiscovery, and the full text is readable by the person with admin rights. Google is the same: their Vault can search, export, and legal-hold Gemini conversations, and if you delete a conversation under hold, “the data is hidden from the user but remains fully retained and visible to Vault administrators.” You won’t see it, but they still can.
Also, these platform vendors allow admins to compare you against other employees. To randomly pick one, Microsoft’s admin usage report shows per-person prompt counts, active days of use, and last-activity per app. This is anonymized by default, but a simple admin toggle will show real names.
That’s just the beginning - there’s a secondary market for this data. Engineering-analytics platforms show quite fine-grained AI use and effectiveness for each engineer, for instance. Cursor includes a “Usage Leaderboard” right in the product. Teramind, an analytics vendor, claims that they “record full ChatGPT sessions showing prompts and responses” and “monitor total messages and frequency patterns per employee.” Pre-AI digital monitoring firms began adding AI-detection modules this spring, ten to nineteen dollars a seat. And ModelOp’s (a governance vendor) survey of a hundred enterprise AI leaders showed that, “almost every Fortune 500 is tracking overall AI usage.”
There’s no meaningful legal constraint here. A few US states require a notification email - the EU goes much further. Surveillance was just included as a default feature suite across the AI stack, and where platforms stop, startups step in.
If you’re a rank-and-file employee, you don’t have anything to show for this. You don’t get insight from this data. You don’t get help on using AI. You can’t prove anything new to anyone. Your career doesn’t get enhanced. You’re just the data donor, and are now subject to a new monitoring, compliance and performance management suite.
If you’re a CTO - or, increasingly, CEO - you may have already extrapolated from the previous paragraph: all this watching is a significant brake and hard limit on the AI transformation that the board is demanding. A key aside: as with my last post, caveat emptor: I cofounded a company that builds what this piece ends up arguing for. The best way I know to counter that is giving you tools to interrogate my claims without trusting me.
Some of you might, at this point, be ready to grab a pitchfork or a torch if I started chanting “Bossware must die!” I agree, the status quo isn’t sustainable, but I won’t call for destruction, exactly. That’s because:
Surveillance is civilizational bedrock
Likely before the clay dried on Sumerian accounting tablets - for labor, barley and beer rations, owed and paid - we have been creating data on work performance to improve quality through standardization and control. More people got fed as a result. Likely many more. 4,000 years later, this sped up: in 1086 William the Conqueror surveyed England so thoroughly that "not one yard of land, nor indeed one ox nor one cow, nor one pig was there left out," and the English named that ledger after Judgment Day. Once per conquest you were asked what you owned. Next uptick: 1888, when a machine stamped a card as you walked through the factory gate. International Time Recording was incorporated as a result and was later merged into the company that became IBM. Twice a day you were asked when you walked through the door. Then, famously, Taylor and his stopwatch, 1911: how fast are your hands, checked per motion. But also Ford's Sociological Department, 1914: inspectors at your kitchen table checking cleanliness and savings before you got the five-dollar day - how do you live, asked at home. By 1948 industrial engineers had published tables pricing every human movement in units of thirty-six milliseconds. By 1987, Congress had commissioned a report on keystroke quotas. By 2020, firms used webcams to watch you work from home. And as of at least 2025, we are basically asking the ultimate question, prompt-by-prompt: what are you thinking?
Intensification aside, all of those 5,000 years have been anchored in four assumptions:
Measurement flows to the watcher. They ask the questions, read analysis, keep the record.
Results run through individuals. Numbers attach to specific people, consequences to numbers.
Goals come from contrast. The bar is set by comparing your people, and the best become the standard.
Clear work benefits the worker. The worker’s return for measurement is constraint that improves their productivity.
Sumerian accounting ran on these. Your monitoring stack runs on them. The sampling rate started at once a harvest and has landed near the pace of thought. And it’s persisted because surveillance has been valuable. By most existential measures life is better than it was even 80 years ago. Safer cars, the cold chain that preserves vaccines, getting paid on time… we have to coordinate to maintain civilization, and watching work closely and honestly is how we solved key parts of this problem. Flat rejection of surveillance would have been close to rejecting progress.
In particular, corporate surveillance of user behavior is clearly valuable. Security teams couldn’t catch exfiltration without it. Remember that OpenAI/HuggingFace breach in July? It was caught and reconstructed via logs. Regulators require logging in a number of industries for audit and trust purposes. You have to pay there, too. And the systems that stop a knowledge worker from pasting your social security number into a consumer chatbot are well worth the money. We can’t have that without reading employees’ messages.
…yet the surveillance bedrock is cracking
There is a yawning, AI-shaped crack through that 5,000-year-old bedrock now - big enough to swallow the transformation your board is demanding. Why is it opening? What are the three ways it’s expanding? Why isn’t it patchable? Read on.
At Meridian Mutual, the prototypical financial services firm I considered in my last post, let’s say pre-AI surveillance might be worth $8 million a year in risk and harm reduction. Perhaps $30 million if there’s a lot of threat in the air. And executives there know the costs: gamed metrics, compliance theater, loss of trust, and the mild crackle of adversarial vibes between the watchers and the watched. Like the rest of humanity, for over 5,000 years, they shrug. It’s all survivable, distasteful friction in an organization that’s slowish to change, where work is legible, and employees were needed. In 2019, a least-worst solution for messy, real business.
Let’s say that taken-for-granted friction - the whole unbilled mess - cost ten million a year. This is down to things like depressed initiative for suggestions, five minutes a week of extra time to look busy, the occasional person quitting as a result. Ten million downside, eight million of real value is a wash and nobody audits those. If it’s thirty million, then you might even relax. For 5,000 years this situation was economically intelligent.
Mass-availability AI has radically upset the economics of all this. Here’s how it shows up on the itemized surveillance bill for Meridian Mutual, the same prototypical financial services firm I considered in my last post:
The tinkering your dashboard suppresses - or drives underground? Think back on my last substack, and on Ethan Mollick’s piece, too: this has gone from an intermittent, low-value suggestion box source to frontier R&D for how to build and run your company. This is the very hour the invoice in my prior post told you to buy for $48 million, because finding out what AI can do for your firm is now the highest-value work in it. So this is an old cost line, repriced by an order of magnitude.
Beyond that, employees can now burn tokens for dashboard theater. It’s trivial for them to ask AI to create fake or token (in the other sense) pull requests, memos, emails. They can even automate this, asking AI to create and optimize software that manufactures this content automatically. Volume is cheap, and if all you’re counting is tokens burned, you’re inviting people to game those boards with your cash. Kings in Sumer did not have to deal with this. New cash, tossed onto the pyre.
Finally, and relatedly: you are now navigating this transition based on surveillance data. For 5,000 years metrics were mostly targets, and targets tolerate gaming - if the number's job is to make people behave, then it’s mostly efficient and at least not deeply destructive for people to manipulate it. But the day AI use became a board question, surveillance data became integral to navigation. Your AI steering committee now reads adoption dashboards to decide which licenses to cut, which groups get to hire, and what the board hears about the transformation - just as gaming went token-speed.
The more you surveil, the more your navigation data is tainted. Censored. Slop-laden. Yet it’s your compass needle as you do AI transformation - redesigning your strategy, KPIs, ops, talent management, you name it. If you ran the organizational transformation harness from my last post, you’d have a self-driving organization but bad GPS. Kings in Sumer could check their tablets to see who was slacking, so this was never a cost line - from 3,000 bc until 2023 ad, the navigation signal was delivered through very human management.
Back in Training is dead I told you your employees are rewriting the org out from underneath you - chat by chat, agentic run by agentic run. That has been running at AI speed since about 2024. It's distributed, chaotic, crowd-driven, and not well documented. So, reminder: there are two transformations underway: the one you're planning and the one that's happening. Surveillance doesn't just corrupt your read on the first. It drives the second into the shadows - the workers hiding their AI use are the people conducting it. You can’t keep up with your own employees while paying to make them invisible.
Skeptic’s Interlude
“Wait a minute!” some of you are thinking. I myself was somewhat skeptical pulling together this invoice and this argument, so I tried to be conservative. As for the substance, there’s more data here than you might expect.
Look at the dials before you argue. KPMG’s survey - one of the best global surveys we have - says 57% of employees already hide their AI use and pass off AI work as their own. Is that more or less than the one hour per week on the invoice in your company? Another pre-AI survey says 46% of tech workers would quit over keystroke tracking alone; above I ratcheted that down to one point of turnover in a hundred. A quarter of workers say they’d surrender real salary to escape monitoring; I settled on one week of hiring friction. And the experimental literature says monitoring makes people break rules more, not less - available studies suggest that watched employees cheat and slow down more, because watching erodes their sense of agency. These popular survey findings match what I see in the field. Over the past two years I've run AI work-redesign sessions at over 40 companies - perhaps 2,500 people, hands on their real work. At least once, in most sessions, someone mentions “big brother”, keystroke logging, AI model companies training on user data. The sentiment isn’t positive, and I’ve never seen anyone contradict that story. And I’ve led a group interview study of software engineers across over 20 firms. There, it was clear that people took sensitive or risky work onto personal accounts and machines, or their phones.
It’s key to note that no dollar is counted twice - the hour lost to hiding isn’t the hour of token performance art, and so on. Basically I set each line item near zero, and the total still lands within a rounding error of the $94.8 million that made our CFO write fk on a napkin in my last post.
Finally, if you think you don’t run bossware: check your SaaS bill. Watching ships as default now - admin dashboards, audit logs, retention policies are most likely available to you. Do your employees believe you aren’t looking at that data? If you’re not sure about that, then the unbilled column is real cost and risk, now. Meridian Mutual is paying $93 million to watch AI use in a way that blocks AI transformation.
Other key objections and responses implied or evident above:
Surveillance has lasted 5,000 years, it’s going to survive! It survived in the results layer. AI is a knowledge work execution engine; watching process taints results.
Costs are higher, fine - we’ll pay. Okay, but you’re paying to corrupt your navigation signal. Worth it?
We’ll fix the metrics. Your people have the same AI you do. Gaming runs at AI speed; fixes run at meat speed.
They’ll put up with it, they always have. Your top talent is the risk you’re taking - they’re most aware and mobile. Ready to roll there?
We'll just watch the agents instead. As of now, this is the same thing as watching people. You’ll see responses to human commands. And if you try to identify those… you’re back in Sumerian territory.
This is all just theory - get out of your ivory tower. Duolingo just reversed course on this. One Very Notable Bigtech Company is burning itself down by doubling down on monitoring.
No matter how much of this invoice we revise, none of these objections eliminate the key bind: watching corrupts the signal, and you need better signal than ever to navigate through this storm.
Our best-in-class answers remain one-sided
Very Smart People have seen all this coming, and there are three serious solution types available today. They improve on the status quo in different ways… and have issues at the same spot.
Club one says: stop watching. Hire adults, measure results, ignore process. GitLab runs its now well-known, all-remote handbook this way - impact over activity, results over hours. Warms my heart as a stance on trust. I have talked to GitLab engineers. Warms theirs, too. As an answer to this piece, however, it’s total restraint: deleting the hidden surveillance bill by taping over the instruments, knowing full well your firm is heading into the most nav-data hungry transition since electrification. This is what pilots might call IFR territory (instrument flight rules, meaning your meatware can’t help you sense the environment), the opposite of what you should be doing. Flying blind with a clean conscience is still flying blind. To be clear, this might win in some cases! Talent is well aware of your choices…
Club two says watch, but blur. Aggregate, de-identify, report at the team level and above. This is the respectable enterprise position, and the platform vendors will sell it to you - Microsoft’s analytics suite promises analysis on “aggregated and de-identified metrics” and even offers differential privacy (e.g., introducing statistical noise to block analysts from naming people). I take this as sincere work by good engineers. But… remember the report at the top of this piece: anonymized by default, but real names one admin toggle away. Anonymity that lives in a setting is a policy, not an architecture, and policies flip under pressure - a bad quarter, a lawsuit, a new CHRO. Your talent - both current and prospective - takes that into account. The blur is binding until the day it isn’t.
Club three - the strongest in my view - says measure better. It began in developer-productivity research: DORA, SPACE, DevEx - Nicole Forsgren and colleagues, the best empirical work ever done on measuring knowledge work, shouting the first half of this piece for years: activity counts lie, individual rankings corrode teams, ask the humans. When McKinsey consolidated that conversation into a framework for scoring individual developers, this turned into a hot, public conflict: Kent Beck and Gergely Orosz pointed out that a favorite executive use for the number is deciding which engineers to fire; Dan North answered with "The Worst Programmer I Know" - Tim, individual score exactly zero because he spent every day pairing, teams that outperformed everything around them. That war is my third question. And this club builds software: DX, Jellyfish, Swarmia and the like run continuous, AI-native telemetry on real work now - they're the engineering-analytics platforms from the top of this piece. Jellyfish reveals itself in its own font, on the homepage: “Know who's using AI, where, how, and which tool.” One sided, and overt about it.
Three clubs. One stops the watching, one blurs it, one sharpens it. Not one changes who the measurement serves - and not one has an answer to the question that decides whether the data comes out true: what does the worker get?
It’s time for mutual telemetry
So - default surveillance is going to beat your AI transformation strategy to the breakfast table.
Flying blind isn’t an option either. I’m proposing an alternative - one that’s only available now because we have AI. Let’s call it “mutual telemetry”: the two-sided measurement of work. The thing that does have to die is the oldest assumption, sun-baked into those Sumerian tablets: measurement flows to the watcher.
Instead, we need to measure more. Far more than any bossware vendor would have dreamed, or could tolerate. My last post runs on this new assumption. The data is the same - turn-by-turn interaction data between workers and models. But we must now provide a two-sided product: private insight for the person doing the work, anonymized, aggregate truth for the firm designing it. Not worker-first charity that compromises the business. Not boss-first extraction. Simultaneous, and valuable for both - because the only way to get honest telemetry is to build a system the watched have a strong reason to feed.
Are you a leader? On a board? An investor? To get great ROI and drive real AI transformation, you have to change the surveillance target, and share the gains with the worker.
Here’s the new deal. Your people keep their AI privacy. No lookup, no rankings, no employer-owned record - and in exchange you get the one thing your prior approach could never deliver: accurate steering data. You give up half of the ranking and rating apparatus and receive a working compass for AI-native organizational transformation. Why half? The top-half of your performance distribution comes for free here. You read that right. Why? People will surface their best work, their adopted skills, the mentoring they do, because visibility finally pays off. People want to know their greatest strengths are valued, and have an impact. Mutual telemetry can prove that, publicly - when the worker chooses that.
You lose the ability to catch the bottom decile shirking. In exchange you get to put your hands on the wheel of a self-driving organization for a nine-figure transformation. Everyone else will be stuck with adaptive cruise control with bad sensor data. Steering data is worth much more than surveillance data. You haven’t had to choose before - we were all in a global trance that there was no other way. Now there is, and it’s mandatory for well-run firms.
One fine Tuesday, when this is all working
Take a look around Meridian’s Tuesday morning in this world. The claims-intake redesign Meridian’s map proposed two months early is up and running, and a senior claims analyst is working within it. She opens her private insights window. It shows six weeks of her performance slope: her escalation judgment is measurably sharper with the new model in the loop; her documentation quality dipped when the workflow changed and recovered in nine days. One thing surprises her: the workaround she built for the most complex class of claims - the ones she used to route around - is the fastest-improving part of her work.
She tries harder things than she used to. Why wouldn’t she? A failed attempt lands in her private record as evidence she’s pushing her edge, not on a manager’s dashboard as evidence for a list. When the workaround delivers, she chooses to share it. Three weeks later four adjusters have adopted it, her name is on it, and she’s the one who opens the staffing conversation about the new escalation pod - holding proof, not a self-assessment.
Down the hall, same afternoon, the CEO opens the aggregate view. No involuntary names in it. No way to put them in. What it holds instead is the thing no surveillance stack ever delivered: where capability is actually rising, who has proudly named themselves and their best-in-class techniques, stats on how frequently their innovations have been adopted, which of those redesigns actually paid off, what the newest model release actually changed about the work - all reported by people whose only way to game the numbers is to actually get better - at learning and performing. The map from my last post proposes; this data decides what’s true. The board asks where’s my ROI, and the answer has line items again - accurate ones this time.
Same turn-by-turn data bossware reads. Opposite physics, because of who it serves. I promised a ruler at the end of my last piece. The worker’s holding it, along with the firm.
Since April of 2024, Juho Kim and I have been building SkillBench, a platform that enables AI native organizational transformation through mutual telemetry. Our early customers can see this future in prelaunch form, now.
But I want to put us aside and give you three questions to ask so you can find out if you have a two-sided solution on your hands.
Before you start, do a zeroth-question “mic check”: does your solution read real work, and not just proxies? These days that’s marrying user chat logs with AI (across all available models) and changes to their work product over time. Model companies can’t do this, but a growing number of traditional surveillance firms do this, so it’s not sufficient. In fact the darkest bossware in history awaits if you stop here.
So, on to the three questions. Two are about the watcher:
Can a manager look up an individual?
Can the system produce a bottom-half list?
And one is about the watched:
What does the worker get?
Your system is bossware unless you get two nos and a third answer that workers love and trust. If you’re thinking you’ve seen this list, you’re right - it’s the bedrock at the top of the piece, inverted: the two nos negate the old assumptions, and the third answer transcends the last one.
On to the new trillion-dollar question: what does the worker get?
A system like this lives or dies based on whether people want it. And for that to happen, the worker has to get real dopamine and a brighter career by using it. The dopamine comes from seeing yourself and others get better, as you connect. The career value comes from being able to prove it.
Remember, this isn’t ethics theater. I’m not here hat in hand for the worker. I’m here saying that if we don’t build mutual telemetry systems that deliver two-sided gains for firm and worker, we will not get the data we need to reconfigure firms at the speed of AI so they can compete in this new age. If we remove the lookup, the ranking, the employer-owned record of employee capability, then the watched have no reason left to shape the data - and huge reasons to feed it. If they start doing that, then the compass works.
You would know a mutual telemetry system was working if workers started trying hard, valuable things without a mandate. Under surveillance a failed attempt gets scored against you. People who leave take their record with them and can show it to get their next job. It will make talent more recognizable in the market. A surveillance record, on the other hand, is a liability you flee and are happy to forget. Candidates will start asking whether you run a two-sided system - just invert the screening question that top talent asks today and you can see the market here. And above all else: people will strongly prefer to be measured.
Run these questions against your surveillance stack. Against any vendor. Against us. Bossware vendors and solutions will be forced to say: that’s impossible for us. Assess the cost of a solution that fails them.
But also, now it’s time to help you falsify my claims.
One. Today, no incumbent monitoring product can answer the three questions correctly. As in - none. Any vendor can falsify with this demo: there is no lookup, it’s not possible to show the bottom-half, and the worker gets insight that they value, personally.
Two. Within a year, the surveillance category will adopt the language of two-sidedness without changing a single answer on number one. The terms “mutual”, “employee-first”, “privacy-forward” and comparable will be on the website, and a lookup will still be in the product. There’s precedent for this: when Microsoft’s Productivity Score shipped per-person tracking in 2020, it was lambasted. But… the fix was removing names from the display. The collection stayed. Expect that move, category-wide, and grade it with the three questions.
Three. "The Duolingo gambit" gets played, on repeat. The namesake: Duolingo wired AI usage into performance reviews, got massive employee backlash, took it seriously, and walked it all back. Before I share the next post in this series (say a month), more name-brand firms will run the same play into the same wall - their own people - and walk it back too. Each time, a firm is discovering the bill above, one line at a time.
The tide is turning, and we must choose
Anti-AI sentiment - in my view - is just barely starting to pick up steam. Executives, shareholders, board members, workers, citizens, everyone is starting to realize the needless harm, waste, and failure associated with AI, and a surprising share of that anti-AI sentiment is anti-surveillance sentiment wearing AI clothing. Every firm that plays the Duolingo gambit finds that out the hard way. What we've lacked is an alternative. The bill above - if turned on its head - is the specification for the solution. You now know how to grade these solutions - SkillBench's included.
Now go further. Think of our shared future. Imagine what happens when a thousand, ten thousand, a hundred thousand, one million, then one billion workers love and trust a substantively neutral work measurement tool for their productivity and career. When firms get the dynamic navigation capability they need to remain relevant and serve the world. That’s worth at least three zeros more than the price of admission. But there’s more here. Workers’ validated capability starts pooling into something bigger and more wonderful than any one firm. Than any occupation. Than any AI model or industry.
On the other hand, be warned: our default approach to AI is stultifying individual and organizational adaptability right when we need it most. If we let that run, things will get… darker. Quickly.
Coarsely, we get the same set of instruments, but we get two different worlds depending on how we handle them. I’ll paint this picture in the last - hopeful and solution-centric - post in this series.
ps: but if you want a dash more motivation to stay tuned, Bill Gates’s note today should give you a sneak peek at the dark side of things. He’s asking for solutions like the one we’ve got in hand.






