10 Comments
User's avatar
Yuki Shoji's avatar

What this makes me wonder is whether AI adoption is often being measured at the wrong unit entirely.

Prompt counts, active days, and tool usage tell us how much AI is being used. They do not tell us whether the work actually became lighter for the person doing it.

If the point of delegation is to remove recurring work from someone’s plate, perhaps one of the most useful signals is not “how much AI did they use?” but “what work no longer required their attention — and what judgement still had to come back to them?”

That seems like a very different kind of telemetry.

Yetvart Artinyan's avatar

Thank you, Matt

“Mutual telemetry” made me wonder whether there is a paradox hidden inside the mechanism itself.

If workers control what becomes visible, that autonomy may be exactly what makes the system trustworthy enough to participate in. But it also means the organization is learning from a self-selected sample of what people are willing to reveal.

The missing data may then be the most valuable data: failed experiments, abandoned AI use, workarounds people would rather not expose, and cases where AI quietly made the work worse.

So here is what I cannot resolve:

Can mutual telemetry ever become representative enough to guide consequential organizational decisions without introducing mechanisms that make it feel like surveillance again?

In other words, could there be an unavoidable trade-off here: the more decision-grade the evidence becomes, the more you undermine the condition that made people willing to generate it in the first place?

What evidence or design principle would convince you that this trade-off can actually be escaped?

Matt Beane's avatar

There are ways to handle this, but in the end there has to be a tradeoff. But there always was. Full-spectrum, constant observation of performance is only accepted when the costs of failure are extreme and failure is irreversible. Surgery doesn't even count there - we don't record video from the surgeon's head! Commercial aviation, nuclear power plants, some policing, etc.

Anyway, the main question is whether it is pareto optimal (for the firm and the worker) to obscure the bottom of the performance distribution in exchange for clear top-line (and anonymized, aggregated, full-spectrum) steering data. My thesis is that this is a good initializing settlement, and we can learn our way forward from there.

Yetvart Artinyan's avatar

Thanks for elaborating. When you say “there are ways to handle this,” which approaches do you have in mind? And what would a Pareto-optimal settlement look like in practice—what could each side see, what protections would each retain, and how would you assess whether the arrangement works for both?

Matt Beane's avatar

I'd rather answer that set of questions through demonstration (aka product).

Hilda S.Nemeth's avatar

I hadn’t really thought about this before and my first reaction to mutual telemetry was actually: I’d love this. I’m exactly the kind of person who would look at my own dashboard and find it motivating.

But by the end I started wondering about the culture it could create. Even without manager lookup or rankings, could voluntary visibility eventually turn into a kind of performance race? If the most ambitious people are sharing evidence of improvement, innovations adopted, etc., does opting not to share really stay neutral?

And perhaps there’s a self-surveillance question too. I can absolutely imagine enjoying my performance slope — and then eventually feeling pressure to keep making the line go up.

Curious how you think about that risk.

Matt Beane's avatar

I so appreciate the thoughtful engagement! You've highlighted a tension in the piece that I left alone.

The first face of what you said has a name in economics: disclosure unraveling. When sharing is voluntary and verifiable, silence starts to seem less neutral. We see this on GitHub's contribution graph today, actually - people make midnight filler commits. The key here - we had thought of this one - visibility attaches to outcomes, never to acts of sharing. A technique travels with your name when others adopt it, but there's no feed of who shared, no streaks, and no way to ask who hasn't. This is not perfect - open source runs on contribution credit and still burns people out. But it's a healthier equilibrium.

Your second point highlights the tension and is also backed by research (people who track their steps walk more and enjoy walking less). The whole point is people attempting harder things, and harder attempts make your short-run line go down! I make this point in my book, too. A naive slope display would punish exactly the behavior mutual telemetry needs to encourage. So the slope has to understand things like difficulty and plateaus, and the tool can't be engagement-farmed.

I sent a note about both of these to my design teams because of this comment. I'm curious: in your own work, what would make sharing feel like contribution rather than subtle entry into a race?

Hilda S.Nemeth's avatar

For me, I suspect mutual telemetry would actually be a net positive. I know myself: give me a metric and eventually I’ll try to improve it 😄 And especially if it also contributed to collective learning, I’d probably enjoy the whole thing (for example, if people also shared things like “I tried this, it failed, and here’s why”).

But I think whether it feels like contribution or a race would depend a lot on the culture and the rules around the system, not just the dashboard itself.

The thing I’d worry most about is context and comparability. I worked in a 120+ person consulting firm where the work across divisions was wildly different. Even two very similar implementation projects could be completely different depending on the client: one cooperative and well-organised, the other full of changing expectations and drama. So I’d be very wary of anything that made those numbers comparable across people without all that context.

And the second one - I realize your design already limits individual visibility for managers. But I think even a rumour that telemetry might eventually affect performance reviews, salaries or promotions could damage trust pretty quickly. There would need to be very transparent communication about what the data will never be used for.

And then there’s the private side. I can imagine my own dashboard being incredibly motivating… until I have a health issue, family problems or just a bad few months and watch my slope go down. At that point the same feedback loop could become pretty demoralising.

So maybe for me the line is: does the system help me learn, or does it gradually make me feel like I should always be improving?

Matt Beane's avatar

Thank you. How interesting and helpful! You... will see your takes reflected in our product suite when it comes out for air (soon). And if we shadows have offended...

Three months ago I decided on an arc for my four "I'm back" substack posts, and post four's target was always your culture-and-rules point. So I hope and expect it will resonate as well.

Latent Dynamics's avatar

The debate over whether workers will hide their failures under a voluntary visibility regime is merely a degraded projection of organizational trust failing its linear identifiability constraints. We've spent millennia treating work telemetry as a behavioral reporting game, asking how much we should show to the watcher. It's a false choice.

The tension disappears the moment we realize that visibility is a massive, unnecessary tax. If we compile work telemetry down to local, spatial registers directly on the worker's device, we eliminate the need for centralized tracking. The worker doesn't need to report. The system doesn't need to look.

Instead of a dashboard, we compile the output into an immutable, cryptographically signed Claim-Level Auditability Graph. This is the third path. Zero-knowledge hardware enclaves verify the soundness of the work and the provenance of the innovation without ever exposing the raw process of the worker. It prevents the voluntary performance race because there are no rankings or visible paths to game. You don't optimize for the gaze of an administrator. You compile against a deterministic verification gate.

The truth is simple: epistemic validation of knowledge work is not a statistical function of behavioral outputs, but an invariant compiled directly into local, hardware-isolated execution planes. When we anchor proof to the substrate, the observer disappears.

If the data is verified locally at the edge, the self-selection bias collapses. We don't need to see the failed experiments to know the verified output is correct. The organization gets bulletproof, decision-grade steering data, and the worker gets absolute, hardware-enforced privacy. The bedrock of corporate surveillance isn't cracking. It's being replaced by something far harder.

If your transformation strategy still relies on someone clicking "share" to prove their worth, you aren't building a self-driving organization. You're just funding a high-speed token performance art gallery. How long until your top talent builds an automated local agent specifically to fake the telemetry? 😉

(⌬_⌬)