Good Protocols · SPC

Writing · 2026-08-26 · Doctrine

Knowing When Not to Measure

A working doctrine for human-serving systems in the AI era

Good Protocols SPC · White paper · Working draft, August 2026


1. The claim, plainly

This paper describes a way of running human-serving work, nonprofit, governmental, social-service, and educational, that measures on a schedule rather than ambiently, records decisions truthfully rather than flatteringly, and uses artificial intelligence to widen judgment rather than to replace it with counting.

We are aware of how this sounds. The sectors we serve have spent three decades building toward dashboards, real-time data, and continuous optimization, often at the insistence of the people who fund them. We are not proposing less accountability. We are proposing a stricter kind: one where the reasons for every consequential decision are recorded at the moment of decision and never edited afterward; where evaluations are dated, versioned, and repeatable rather than final; and where an organization states in advance, in writing and in its proposals, what it will not measure and why.

The doctrine rests on one narrow claim, and the narrowness is the point: continuous ambient measurement of relational practice degrades the practice it measures. Not all measurement. Not financial controls, not safety reporting, not the disaggregated outcome data that equity work requires. Specifically, the always-on scoring of the human moments where the actual work of these sectors happens: a caseworker with a family, a teacher with a student, a navigator with a newcomer. Those moments do their work partly because they are not fully administered. Instrument them continuously and you get better-looking numbers describing worse work.

For fifty years the honest response to this problem was a shrug, because counting was the only thing that aggregated. Qualitative judgment didn’t scale; numbers did; so numbers won, and everyone learned to live with what Campbell warned about in 1976. That constraint has now lifted. Language models can read, compare, and aggregate structured qualitative judgment at scale. The historic excuse for converting practice into counts is gone. What remains is the question of discipline, when to look, what to record, and what to leave alone, and that question is what this paper answers.

2. Why: what the dashboard actually costs

The case against ambient measurement is not romantic. It is one of the better-evidenced findings in the social sciences, made by people who loved data.

Donald Campbell observed that the more any quantitative indicator is used for consequential decisions, the more pressure it comes under, and the more it corrupts both itself and the process it was meant to monitor. Charles Goodhart made the same observation about monetary targets; it now carries his name. W. Edwards Deming, the patron saint of measurement and not its enemy, insisted that the most important figures for running an organization are unknown and unknowable, and that managing only by the visible numbers is one of the deadly diseases of management. Michael Power documented how the audit explosion of the 1990s made organizations auditable rather than good, redirecting effort from the work to the appearance of the work. Onora O’Neill, in her 2002 Reith Lectures, traced how the machinery built to secure public trust (targets, league tables, perpetual reporting) was corroding both trust and the professional judgment it depended on. James C. Scott showed what happens when states make complex living systems legible enough to administer from above: the map becomes the territory, and the territory dies back to match the map. Jerry Muller gathered the modern case histories under an accurate title, The Tyranny of Metrics.

Anyone who has worked a frontline job in our sectors can supply the local translation. The caseworker who spends more of the week documenting encounters than having them. The teacher whose data wall knows things about her students that she no longer has time to know herself. The shelter worker whose every interaction feeds a system that will one day be used to prove the interaction happened, and never to understand it. Study after study of direct-service work finds documentation burden named among the leading drivers of burnout and exit, which means ambient measurement is not merely failing to capture the work; it is consuming the workforce that does it.

The sociologist Hartmut Rosa gives the deepest account of why. His theory of resonance holds that the moments where human-serving work actually works, where a person is reached and reaches back, require Unverfügbarkeit: a partial uncontrollability, a remainder that stays beyond administration. A moment rendered fully visible, fully scored, and fully steerable cannot answer back; it can only comply. Rosa’s recent work distinguishes the Situation, the lived moment entered from inside, from the Konstellation, the arrangement of forces seen from above. Both views are necessary. The error of dashboard culture is not that it builds the constellation view; it is that it leaves the constellation view running all the time, projected onto the situation while people are trying to live it. The overhead light never turns off, and then we wonder why nothing intimate happens in the room.

None of this is an argument that the sectors’ data-builders were foolish. They were working under a real constraint: funders and legislatures needed aggregate accountability, and numbers were the only artifact that aggregated. The people who built the dashboards were solving the problem honestly with the instruments of their era. The instruments of the era have changed. The doctrine below is what accountability can look like now.

3. What: the doctrine in five principles

Our house keeps this doctrine in the language of the loom, since we are woodworkers and weavers before we are technologists and the craft vocabulary earns its keep, but each principle translates fully into sector terms. Both registers are given.

First: the frame recedes. Constellation instruments, meaning anything that renders the whole arrangement at once (the rubric grid, the systems map, the aggregate view), run at defined stations only: at dressing (planning, before the work begins), at breakdown (when something has failed and diagnosis is needed), and in traffic (scheduled re-readings of past work). Between stations, they are out of view. A weaver studies the loom’s frame when dressing it and when tension fails; while weaving, a good frame disappears. In sector terms: measurement has a calendar, published in advance. There is no ambient mode. No live dashboard sits on the working screen of a person doing relational work.

Second: decisions owe truthfulness, not pattern. The present of any working system is what weavers call the fell, the exact line where the last pass was beaten in and possibility became fact. At that line, what an organization owes its future self is an honest record: the situation as it looked, the options actually considered, every party rendered as a decision-maker with reasons, the choice made, and, critically, the uncertainty and dissent preserved verbatim. What it does not owe is coherence with a pattern, because pattern is only visible from rows away. Demanding that each decision justify itself against a strategy no one can yet see produces confident fiction. Recording reasons as they were, sweat included, produces an archive worth inheriting.

Third: the past is closed in fact, open in traffic. The record never changes; what the record means changes constantly, because every new event re-reads the old ones. Three red rows are a stripe until the eighth pass makes them plaid. So evaluations are dated, versioned readings, kept beside every earlier reading, never overwriting them and never final. Re-reading is legitimate work, not revisionism. One warning travels with this principle: records lie flat. The archive smooths every frantic week into apparent calm, which is how institutions launder panic into inevitability and then draw the wrong lessons. Honest re-reading includes restoring the texture the record smoothed away.

Fourth: improvisation is parasitic on fixity. The objection we hear most is that this much record-keeping will deaden the responsive, improvised character of good frontline work. The truth runs the other way. Musicians are free precisely because the changes don’t change; the ground bass is what makes the variations possible. Practitioners improvise best over a stable, honest archive, knowing what was actually decided and why, and worst under ambient scoring, where every move is judged mid-motion. Fixity of the record is the floor of freedom, not its ceiling.

Fifth: sampler and cloth. Every working space runs in two modes, and the boundary between them is enforced technically, not promised politely. Sampler is the weaver’s side piece, the mode for trying postures, running counterfactuals, arguing, and riffing, with every instrument available and nothing retained. Cloth is the record, entered only by a deliberate act, at which point the decision and its reasons are written, append-only, into the archive. Nothing reaches the cloth untried in the sampler; nothing lingers in the sampler once the room has decided. In sector terms: exploration is unlogged by construction, and decisions are logged completely. Most current systems get this exactly backwards: surveilling the exploration and losing the reasons.

4. How: the instruments

The doctrine is implemented, not aspired to. Four instruments carry it.

The situation rubric. Evaluation in this house is rubric-based evaluative reasoning in the tradition of Jane Davidson: explicit criteria, defined levels, evidence marshaled into a transparent judgment that another reader can follow and contest. Our criteria come from Jens Beljan’s account of what formative environments must sustain: Anschlussfähigkeit (connectability: can others tie onto what we do?), Beweglichkeit (mobility: does this move keep future moves open?), and Erweiterung (expansion: does the fabric widen, with new participants and new edges?). Each criterion is read at four levels, the child, the worker, the family, and the community anchor, yielding a twelve-cell reading of any candidate course of action. The rubric runs only at the three stations. At planning, it reads options before commitment. At breakdown, it diagnoses. In traffic, it re-reads past decisions as their meaning shifts. It is never a dashboard.

The two ledgers. Every consequential decision produces a fell record, written at the moment: timestamp, the situation as described, the parties each rendered with their reasons, the postures considered, the one chosen, the stated rationale, and the uncertainties and dissent, verbatim. It is append-only and is never edited by anything that comes later. Separately, a readings ledger accumulates the dated rubric readings made from rows away. The two never merge. Reasons live in one place, judgments in the other, and a funder, board member, or auditor can trace both. We would put this archive up against any dashboard as an accountability instrument: it shows not just what happened but what we believed and doubted while deciding, which is the thing dashboards structurally cannot show.

Declared unmilled zones. Every tool we build, and every proposal we submit, states in writing what it will not measure, surface, or automate, as a feature of the specification rather than an apology. A woodworker leaves the live edge; the tree drew that line first. Declaring the unmilled zone in advance converts a suspicious silence (“what aren’t they tracking?”) into a professional judgment a funder can evaluate and accept. It is more honest than the industry norm, which is to imply comprehensive measurement and quietly deliver partial.

The AI, deployed as scout rather than advocate. This is where the era actually changes. Language models make it possible to do at scale what only counting could do before: aggregate judgment. In our rooms the AI has four standing functions. It refuses scenery: before any evaluation, every party in a situation is rendered as a decision-maker with reasons, because a system that lets any stakeholder become backdrop will optimize against them. It carries the counterfactual fan: for each candidate course, it retrieves comparable past decisions and their outcomes, under a standing, non-disableable instruction to lead with the precedents that disfavor the option the room walked in preferring, which is the discipline no group of humans under pressure reliably keeps for itself. It reads the twelve cells on request, at the stations only. And it logs the pass faithfully: reasons as given, not as improved. Note what is absent: the AI does not score people, does not monitor practice ambiently, and does not predict individuals. It widens the judgment of the humans at the fell and keeps their archive honest. That is the whole job.

5. What we keep

The doctrine governs the measurement of practice. It does not touch, and we affirmatively keep, the controls that protect people and public money: financial accounting and audit, incident and safety reporting, mandated child- and adult-protection reporting, statutory data submissions, and the disaggregated demographic and outcome data without which no organization can see its own disparities. Equity analysis in particular is station work we schedule and take seriously; the doctrine’s target is the ambient surveillance of workers and clients, which has never been what catches inequity. Scheduled, disaggregated re-reading is. Where a government data system (an HMIS, a state student-information system) requires submissions, we comply, and where possible we generate the submission from our ledgers rather than building the practice around the submission. The discipline cuts both ways: we refuse ambient measurement, and we refuse the lazier temptation of refusing measurement altogether.

6. The development problem, named

We will not pretend this ethos is convenient to fund. It makes the conventional development role, the person whose job is to promise dashboards, produce them, and mine a donor database in between, genuinely difficult to define. We think the difficulty is the beginning of a better definition, and since we intend to hire and keep such a person, we owe the definition in writing.

What the role stops being. It stops being a metrics vendor. It does not promise real-time impact dashboards, does not invent proxy numbers to decorate proposals, and does not convert program staff into data-entry labor for reporting’s sake. It also applies the doctrine to its own craft, and this is the part that surprises people: donor relationships are relational practice, so there is no ambient scoring of donors, no moves-management machinery that annotates every human contact in a pipeline. Gift processing, acknowledgment, and restricted-fund accounting remain, because those are fiduciary controls.

What the role becomes. Three functions. Qualification: finding, and cultivating over years, the funders whose theory of accountability is compatible, and they exist in growing numbers. The trust-based philanthropy movement, the turn toward multi-year unrestricted giving exemplified at the largest scale by MacKenzie Scott’s practice, participatory grantmaking, and the funders influenced by Edgar Villanueva’s critique have all arrived, by their own routes, at the position this paper argues from the practice side. Our development officer’s first market is that alignment. Translation: authoring, for every grant, a measurement schedule, a one-page attachment stating the stations at which we will read, what the fell record captures, what the readings ledger will show the funder and when, and the declared unmilled zone. This replaces the KPI appendix. It gives a program officer something concrete, defensible, and frankly more auditable than a dashboard, and it moves the negotiation to the honest question: when shall we look together, rather than how continuously can we pretend to. Refusal stewardship: maintaining the record of funding we declined and why, written at the moment, reasons unimproved, like every other pass. Over years that ledger becomes the organization’s credibility, the asset that plain dealing has always been in every field that keeps archives.

What the role refuses to imitate. The conventional development function was conscripted into an arms race the sector rarely names: funding processes that preach collaboration and practice competition, where a request for proposals draws five and ten times the plausibly fitted applicants because organizations converge on money mimetically, wanting what the others want and casting nets far outside their strengths to buy breadth, until the convergence manufactures the very scarcity everyone fears, funders drown in misfit proposals, and missions drift toward whatever was last funded. Our development office declines the race, not just particular entries in it. The grain test governs every application: we apply only where the work demonstrably matches our existing strengths, fit we can specify in advance in the measurement schedule itself, and wide-net applications are declined with the reasons recorded, like every other pass. One further practice, kept as policy because no one keeps it as a mood: when funding we sought lands with a better-fitted peer, we honor them in writing. The annual station reading of the development function asks whether we did. An organization that cannot congratulate the winner of a grant it wanted is being formed by the race it claims to have left.

The honest costs. The compatible funding pool is smaller than the general pool, so budgets are built to the aligned money, not to the theoretical market, which means slower growth and real concentration risk, managed by breadth within the aligned pool rather than by dilution of the doctrine. Some government contracts specify ambient monitoring of practice as a condition; those we renegotiate where we can and decline where we cannot, and the decline goes in the ledger. A development professional who joins us trades the larger hunting ground for a defensible one: they will never have to promise a funder something the program side quietly resents delivering, and every claim in their case for support is one the organization can survive an honest audit of. There are fundraisers who have been waiting their whole careers for that trade. The role is difficult to define only against the old job description; against the work, it is the cleanest development job in the sector.

7. Objections, answered briefly

“This is anti-accountability.” It is stricter accountability. Append-only decision records with dissent preserved, dated evaluations that never overwrite, and pre-declared limits are harder to game than any dashboard. What we refuse is the theater of accountability: the continuous feed that measures compliance-shaped behavior and calls it the work.

“Equity requires data.” Yes, and it gets data: disaggregated, scheduled, and read seriously at stations. What equity has never required is the ambient surveillance of frontline workers and the people they serve, which falls heaviest, as surveillance always does, on the least powerful people in the building.

“Our funders and legislators require dashboards.” Some do, and some of those will accept a measurement schedule once they see one, because most program officers privately know what Campbell knew. Some will not, and that is what the refusal ledger is for. An organization that cannot name the money it won’t take doesn’t have an ethos; it has a mood.

“This cannot scale.” Counting scaled because it was the only thing that did. Structured qualitative judgment (rubric readings, fell records, comparative retrieval across an archive) now aggregates, because reading at scale is precisely what language models do. The scale objection was true for fifty years and is not true now.

“You are romanticizing the unmeasured.” The lineage says otherwise. Davidson’s evaluative reasoning is rigorous method; Rosa’s is a systematic sociology; Campbell, Deming, Power, O’Neill, Scott, and Muller are the measurement tradition criticizing itself with evidence. We are not against knowing. We are against pretending that the always-lit dashboard is how knowing happens.

8. The discipline, stated as covenant

Because this ethos will be pressured, by funders, by procurement, and by our own anxiety in lean quarters, we state the lines we hold, so that holding them is policy rather than heroics:

No ambient mode, for anyone, including funders and including ourselves. No invented numbers in any proposal, report, or page we publish. Reasons recorded at the moment of every consequential decision, dissent included, unedited forever after. Every reading dated and versioned; no reading final. An unmilled zone declared in every specification and every proposal. Declines recorded with reasons, like every other pass. And tone discipline: no contempt for the colleagues and public servants who built the dashboard era; they solved the aggregation problem honestly with the instruments they had, and we can hold our doctrine without sneering at the people whose constraint has only just lifted.

One last reading, of this paper by its own rubric. Connectability: it ties onto living movements, trust-based philanthropy, the evaluation field’s own rubric tradition, and the measurement literature’s self-critique, so others can tie onto it. Mobility: it commits us to a schedule and a record, not to any frozen program design; every future move stays open. Expansion: the instruments are built to be adopted piecemeal, a single measurement-schedule attachment or a single fell record, so the fabric can widen one organization at a time. The doctrine passes its own reading. Now it goes to the loom.


Selected lineage

Donald T. Campbell, “Assessing the Impact of Planned Social Change” (1976) · W. Edwards Deming, Out of the Crisis (1986) · Michael Power, The Audit Society (1997) · James C. Scott, Seeing Like a State (1998) · Onora O’Neill, A Question of Trust (Reith Lectures, 2002) · E. Jane Davidson, Evaluation Methodology Basics (2005) · Jens Beljan, Schule als Resonanzraum und Entfremdungszone (2017) · Jerry Z. Muller, The Tyranny of Metrics (2018) · Edgar Villanueva, Decolonizing Wealth (2018) · Hartmut Rosa, Resonance (2019) and The Uncontrollability of the World (2020) · Hartmut Rosa, Situation und Konstellation (Suhrkamp, 2026; English translation in preparation at Good Protocols) · Trust-Based Philanthropy Project (2020– ) · Paul Andrew Hutton, The Apache Wars (2016) and The Undiscovered Country (2025), for the archival warrant: a record in which every party remains a decision-maker is the only archive worth inheriting.

Good Protocols SPC builds AI-era instruments for human-serving systems. Companion documents: foundations.md (the five principles as governing doctrine) and the Steward space build brief.


All writing