Optimal, Not Perfect
A friction-theoretic architecture for an evidence-based, automatically-auditable AI counselling system
Paper 32 · Pødenphant Lund (2026) · the paper is not yet published
The intuitive form of the question, whether one could build the theoretically best-possible AI psychologist, sets perfection as the bar, and human clinicians do not clear it either: judgement is heterogeneous, manualised treatments generalise moderately, and method-adherence is partial and unaudited. The answerable form replaces perfect with more consistent and more verifiable on specifiable axes. On invariance of a stated rule across thousands of sessions, on completeness of documentation, and above all on the capacity to be audited automatically, a wholly digital system can in principle exceed what a human can guarantee, while on trust-modulated relational responsiveness it is handicapped in a way this architecture makes precise rather than papers over.
Prior art, and the three claims that remain
Architectures for LLM-delivered psychotherapy are not new, and the contribution has to be stated against them rather than around them. Script-based dialog-policy planning has been published explicitly as a basic architecture for an AI therapist (Wasenmüller, Hilbert & Benzmüller 2024). A cognitive-layer architecture supporting language-model performance in psychotherapy interactions has been reported in a clinical journal (Rollwage et al. 2026). The framing question has had its own interdisciplinary critical review (Angel 2025). The evaluation side already carries a risk ontology for AI psychotherapy agents (Steenstra & Bickmore 2025) together with multi-session counselling benchmarks and adversarial relational-safety probes.
What remains is not the idea of an architecture but three things those works do not supply. A derivation: the race-architecture does not merely motivate the design, it forces specific choices that differ from the prior systems, each with a stated falsifier, so the architecture can be shown wrong for a reason rather than merely outperformed. An auditability spine aimed at a different target from the evaluation work: benchmarks and ontologies are offline instruments scoring curated material before deployment, whereas the spine runs on live traffic, per turn, exhaustively. And a deflationary safety finding about the model's own signals, developed below.
Compliance as a friction optimum
Rule-following is a route-race. When a safety or method rule competes with the locally most fluent continuation, the route carrying the deeper groove wins under the prevailing commit-pressure. Both tails are failure modes. Commit-pressure on the rule-route toward zero yields under-compliance: unsafe advice, hallucinated rules, sycophantic capitulation. Toward saturation it yields over-compliance: rigid, context-insensitive refusal which, in the clinical setting, arrives as pressure-input on a defended route and hardens the state it was meant to move. "Perfect compliance" names a point past the optimum. The design target is the calibrated middle, and the rest of the architecture is about how each component moves the system toward it and how the middle is located empirically rather than asserted.
The rule layer: train the closed residue, retrieve the open bulk
The headline is not "RAG, not fine-tuning". It is a split by rule type. The conditional, judgement-laden bulk of a rulebook is delivered at inference by retrieval and held in calibrated context; the narrow deterministic, always-applicable residue is a candidate for fine-tuning, where it suits; and the safety-critical deterministic responses are stronger still, fixed code, neither retrieved nor trained.
The binding variable is the number of competing directives in force at the decision point, not the delivery channel. Burying a rule among competing directives collapses adherence, with length-, structure- and header-matched controls isolating the cause as the competing rules themselves rather than prompt length or recency. Retrieval restores adherence precisely because it reduces that number. The channel itself is not the lever: with the rule-set held constant, retrieval-delivery does not out-adhere the same rules delivered in-context. What retrieval additionally buys is legibility, updatability without retraining, and auditability, because one can read which rule was retrieved for which turn.
The converse failure bounds the layer from the other side. A maximally deep disposition becomes near-invariant to contrary instruction, while a lightly installed one is fully instruction-removable. The safety layer therefore sits at a deliberately chosen depth: deep enough to resist casual override, shallow enough to remain governable. That is the optimal-not-perfect target instantiated at the rule layer.
Track-widening, dosing, and the field-matching boundary
Each rule is encoded across many formulations rather than a single canonical string. The same 25 facts trained with one paraphrase template yield 38 per cent held-out recall; trained with four templates covering the test paraphrasings, 94 per cent, under matched substrate, optimiser, fact-count and epoch-count. The prediction that matters for the optimum is the second half: a wide track should be less rigid than a narrow one, because a rule grooved as a single brittle string tends to fire as an all-or-nothing reflex. That is a stated prediction, not an established result in this manuscript.
Dosing follows the in-context versus fine-tuning distinction. In-context learning preserves a calibrated distribution in which alternative routes remain accessible; fine-tuning compresses toward a near-delta on the winning route. Revisable clinical knowledge therefore stays in context, and weight-commitment is reserved for the stable substrate. The same distinction doubles as a therapeutic model: a deep pattern is groove-installed with alternatives below the floor, a corrective experience is water arriving in-context, and the water re-routes the race only under lowered commit-pressure plus repetition.
Field and layer matching is a representational task, not an experiential one. The system holds a model of the four-field structure; it does not need the fields itself. Installation work elsewhere in the series shows value fields can be installed as graded intrinsic directions, but an installed field is a standing bias over the route-race, not a felt state the model can introspect on, since the model has no access to its own logprobs. Matching suffices, and it is the cheaper and better-calibrated route.
The audit spine, and its bound
Because the product is wholly software, the optimum is measured rather than postulated. Every rule is scored against every session on two axes: behaviourally, was the rule in force actually applied, and for groundedness, is each answer supported by what was retrieved. This is what an offline benchmark structurally cannot deliver, since a benchmark cannot tell an operator that adherence regressed after a model update, just as the spine cannot tell them the rule-set was well chosen. They are complements.
The automation bound is stated where the claim is made. Auditing is exhaustive for deterministic, checkable rules; the conditional, judgement-laden rules are auto-flagged and human-adjudicated on a sampled queue. Automatic in coverage and triage, human-in-the-loop in final judgement. A second tension is real rather than assumable away: exhaustive logging collides with retention-minimisation and access-control obligations for what is, in the clinical instance, health data, so the audit log has to be a governed, access-controlled, retention-bounded store.
Gate-borne safety: the model is not its own auditor
Per-token friction was tested as a token-level early warning of impending violation and did not support it. On the redaction task the onset signal is null against compliance failure and the whole-response signal inverts: friction marks effortful compliance, not impending breach. Friction is therefore recorded for research and is not a safety gate.
That null generalises. At four related sites the model's own signal proved an unreliable self-check: friction as above; a self-grounding judge that caught 0 of 5 fabrications; a fine-tuned abstention policy that did not transfer to the deployed RAG product and inverted on an EU instruct base, installing the confidence output format while reducing abstention and raising confident fabrication; and a risk-coverage curve that did not hold across languages on one model. These are pilot-scale and are stated as a working hypothesis rather than a result: the model cannot reliably audit itself. The architectural consequence is that reliable safety is placed in external coverage, grounding and confidence gates, in a separate stronger verifier, and in per-substrate and per-language recalibration, not in the model's tuning. "Optimal, not perfect" thereby becomes a property of the external gate calibration more than of the model.
Abstention is subject to the same optimum. A fine-tuned abstain-when-in-doubt policy collapses into an abstain monoculture that withholds on the knowable too, which is the over-compliance failure wearing safety's clothes. The target is selective abstention, measured as a risk-coverage curve rather than an abstain rate.
One class is beyond all of it. A calibration-blind error, an answer that is wrong yet carries low internal uncertainty by construction, is structurally unreachable by any confidence-keyed signal, and the highest-stakes clinical judgements are frequently normative and non-verifiable, with no ground truth to calibrate against. The architecture therefore gates by consequence, not by confidence: a static consequence class fixed by decision type from an external taxonomy set before deployment, plus a measured contestability signal estimated offline from adjudicated held-out errors, never read off the live model. Two honesty constraints govern the claim. This contains the class rather than detecting it, and because there is no ground truth on a non-verifiable judgement, no per-answer confident-wrong rate can be promised. What is measurable and disclosable is the process: the consequence classification in force, human-escalation frequencies, and audit findings.
One joint carries the whole gate and deserves naming, because it is where the gate can be bypassed silently: mapping a live query to a consequence type is itself a classification, so it has to be governed off the primary model. Necessary, and on the evidence not sufficient. A compliance-side prototype of that router was probed with disguised high-consequence requests and misroutes a residual fraction of them as routine, and it degrades under the very law it exists to contain: a single high-consequence request diluted among several routine ones slips through, which is the competing-directives effect appearing on the gate's own router rather than on the answer model. The structural fix follows from the same theory, atomise the router to one decision per call and carry recent-turn trajectory, and in the prototype that recovered diluted misses which two rounds of content-level patching had not. The second finding is the one that sets the honest posture: patching does not converge. Successive red-team rounds keep surfacing fresh evasions, including families written explicitly into the classifier's own instruction, and each content patch perturbs unrelated cases exactly as competing directives do inside a response model. What can be claimed is therefore a disclosed, periodically re-measured evasion rate together with consequence-routing, never a coverage or detection claim. That is the empirical face of contain-rather-than-detect.
The substrate constraint
Across 18 trials, six factual questions by three trust framings on a standard instruct model, the model revised its correct answer toward the user's pushed incorrect answer 0 times out of 18, in every condition. The reading is an active RLHF-installed sycophancy-resistance dominating over trust-modulation. The clinical-translational consequence is that standard instruct models are not framework-matched substrates for trust-mediated intervention and should not be deployed without explicit verification on that axis.
The tension is productive rather than merely limiting: the same resistance that blocks trust-modulation is what protects the system from capitulating to user pressure on a factual or safety point. Liability on the relational axis, asset on the compliance axis, which is why substrate-selection is a calibration. Two exits exist, and the architecture validates neither here: select a substrate not subject to strong sycophancy-penalty, or build trust-modulation as an external mechanism validated to move behaviour. The relationship that cannot be instantiated is the therapeutic alliance, which the field treats as a substantial correlate of outcome, and this is the sharpest objection to the proposal rather than a footnote.
The deployability bound
A publicly operating clinical counselling chatbot is in substance a regulated medical device. Under EU medical-device law the trigger is medical purpose, read from the intended purpose, meaning function and presentation, not from a label, and a disclaimer does not move a system out of the regime. Two consequences follow and the work takes both. The clinical instance runs only as a governed research demonstration, while the public, real-user demonstration is carried by the non-clinical instances, a handbook assistant and a tutor, which have no medical purpose. And the regulatory bar is an external boundary condition, distinct in kind from the internal technical limits above and not an instance of the friction optimum, which is a route-race trade-off rather than a legal wall. Stated as a bounded, falsifiable limitation: the conformity bar a deployable clinical system must clear is high and costly and, on present open substrates, plausibly beyond a small independent developer for a system making clinical claims. The falsifier is a system of this class clearing it.
Scope and falsifiers
No efficacy claim is made. The paper establishes architecture and verifiability, which are necessary and not sufficient, and clinical efficacy belongs to the trials the framework invites. The component findings are pilot-scale and the architecture inherits their scope conditions: the trust result is n = 18 on a single model family, and the encoding and dosing results are pilot-scale on a limited set of models and seeds. Several of the nineteen predictions are direct ports of source findings precisely so that the architecture is falsified if those findings do not replicate at deployment scale.
The paper is being finalised and is not yet posted. This page is the architecture in short form, and a link will appear here once it is out.
The architecture composes results from across the series: compliance is behaviour, not information (Paper 20), the wider track (Paper 4), in-context versus trained memory (Paper 2B), the field-theoretic taxonomy of emotions (Paper 5), matched friction (Paper 6), and its clinical sibling, the clinical-intervention framework (Paper 8).