Claude

IntuitionLabs is now a member of the Claude Partner Network – AI training and upskilling with Claude for pharma and biotech. Book a call.

IntuitionLabs
Back to Articles
IntuitionLabs

claude api on-demand compaction · claude context compaction

Claude API On-Demand Compaction: Pharma Agent Validation

October 5, 2026
26 min read

A 2026 analysis of Claude API on-demand compaction for pharma agents, covering signed blocks, current platform support, preserved thinking, critical-state regression tests, audit evidence, and measured token economics.

Claude API On-Demand Compaction: Pharma Agent Validation
Summary
  1. 01Treat compaction as a lossy state transformation. Transport acceptance, critical-state preservation, and correct authorized continuation require independent verdicts.
  2. 02Define a critical-state contract covering identifiers, adverse-event meaning, evidence, protocol constraints, temporal updates, pending actions, and unresolved questions.
  3. 03Test uncompacted and compacted branches against independent authoritative labels. Vary evidence position, retained-tail boundaries, corrections, and repeated replacements.
  4. 04Preserve action status outside the summary, enforce authorization downstream, and reconcile pending operations before reissuing work during recovery.
  5. 05The article reports no paid API experiment, savings percentage, or preservation rate. Benchmark research informs test design; economic decisions require actual session usage and category-specific costs.
01

Executive Summary

Claude API on-demand compaction gives developers control over when a conversation is summarized through the application programming interface (API). Anthropic introduced the beta on September 14, 2026; the release described a signed block, background execution, and retention of recent turns. ([1]) This report's central recommendation is to treat summary replacement as a lossy state transformation with explicit acceptance tests. A theoretical compaction paper examines selection and bounded summarization, but models a single invocation; it provides evaluation framing rather than a guarantee for a production agent. ([2]) ([3])

The compact-2026-09-04 header applies to generation and later requests carrying the block. ([4]) Current documentation, checked for this report on October 5, 2026, lists the Claude API, Claude Platform on AWS, Google Cloud, and Microsoft Foundry in beta. Amazon Bedrock is absent from that on-demand support list. ([5]) Kept thinking has additional history and model constraints, so successful summary generation alone is an incomplete acceptance check. ([6]) A signature establishes an API acceptance mechanism; semantic completeness must be evaluated separately against authoritative state and intended use. The applicable record requirements must first be established rather than assumed. ([7])

For pharmaceutical workflows, the proposed golden set covers case identity, adverse-event facts, evidence provenance, protocol constraints, knowledge updates, and pending actions. It tests retrieval and decisions before and after replacement, including repeated compactions and concurrent tool activity. Long-context research supports position-sensitive testing; long-term memory evaluation also covers temporal reasoning, updates, and abstention. These findings guide test design and do not measure this compaction feature. ([8]) ([9]) Regulatory scope remains workflow-specific: the Food and Drug Administration (FDA) recommends documented, risk-based validation, while closed-system controls include accurate and complete record copies. ([10]) ([11])

The economic decision requires actual request usage. Current Claude Sonnet 5.5 list prices are $2 per million input tokens and $10 per million output tokens. ([12]) Token counting provides estimates for planning. ([13]) This report performs no paid API experiment and reports no savings percentage or preservation rate. Adoption should require a passing semantic regression suite, recoverable original history, controlled tool permissions, and an auditable replacement process. These are proposed engineering controls, supported by provenance standards and downstream authorization guidance, rather than vendor or regulatory certification. ([14]) ([15])

17

Long-context models evaluated by RULER, not a pharmaceutical compaction performance result

13

Representative RULER tasks extending retrieval with multi-hop tracing and aggregation

500

Curated LongMemEval questions covering extraction, reasoning, updates, and abstention

02

Introduction and Background

Long-running pharmaceutical agents accumulate more than conversational prose. A useful validation boundary is the state needed to continue the task correctly: the governing protocol, authoritative case facts, cited evidence, unresolved questions, and permissions for future actions. This report proposes making that boundary explicit before implementing context compaction. The rationale is consistent with the World Wide Web Consortium (W3C) definition of provenance, which connects data to the entities, activities, and people involved in producing it. ([14])

Context compaction replaces older conversation turns with a summary. ([16]) It therefore changes the representation supplied to the model. The original transcript may remain in application storage, but the compaction theory paper explains that discarded information no longer reaches the model unless subsequent inputs restore it. This distinction motivates separate controls for record preservation and operational recall. ([17]) A recovery archive is useful only if the application can retrieve the relevant material when needed.

The audience is architects, machine-learning engineers, validation teams, and owners of safety, regulatory, clinical, and scientific agents. Good practice, commonly abbreviated GxP, is used here as a collective label for regulated practice domains; the Medicines and Healthcare products Regulatory Agency (MHRA) guidance covers manufacturing, distribution, and pharmacovigilance among its sectors. ([18]) Applicability must be decided for the actual workflow. FDA's Part 11 guidance recommends documenting those scope decisions. ([7])

IntuitionLabs' adjacent advisory perspective is reflected in its stated emphasis on connecting artificial intelligence (AI) to authoritative enterprise sources with identity, permissions, retrieval, citations, and evaluation. ([19]) The technical analysis below consequently focuses on demonstrable continuity of work. The Organisation for Economic Co-operation and Development (OECD) similarly frames computerized-system validation around risk assessment and a lifecycle approach. ([20]) ([21]) The proposed harness is a concrete implementation design, not a claim that these sources prescribe a particular AI architecture.

The regulatory sources serve different purposes. FDA's scope guidance describes recommendations, while Title 21 of the Code of Federal Regulations (CFR), section 11.10, supplies applicable closed-system controls; OECD guidance emphasizes scalable risk assessment. These should be read according to their scope, not combined into an automatic approval of an agent implementation. ([7]) ([22]) ([20])

F.01
Selected Claude Sonnet 5.5 token prices
03

Key Changes

Developer-controlled timing and platform scope

Threshold compaction runs within the request that reaches its trigger; on-demand compaction lets the developer choose the moment. ([16]) Table 1 compares the operational choices. Its validation column contains proposed checks rather than documented performance promises.

T.01
PatternDocumented behaviorProposed acceptance check
ThresholdSummarization occurs in the threshold-reaching request. ([16])Reproduce the boundary and inspect the next task decision.
On demandDeveloper chooses when to request a signed summary. ([1])Confirm the selected prefix matches the archived snapshot.
BackgroundFull-history work continues while compaction runs. ([23])Reject a result if its snapshot lineage is stale.
Keep recent turnsApplication chooses the cut; tool calls and results stay together. ([24])Compare retained messages exactly and test pending work.

The choice should follow the task's state transitions. This report recommends a boundary after evidence has been committed and before a new decision phase begins. That is an application policy, not a vendor timing requirement. Modeling replacement as derivation makes the old state, transformation, and new state identifiable. ([3]) Atomic transaction mechanisms offer a useful implementation pattern for committing that transition. ([25])

The current supported model families are Fable 5/5.1, Mythos 5/5.1/Preview, Opus 4.6/4.7/4.8/5/5.5, and Sonnet 4.6/5/5.5. ([26]) Teams should pin the selected model and platform in qualification evidence; availability on one named service should not be generalized to another service with a similar cloud label.

Signed-block lifecycle

Replace the summarized prefix with exactly one unchanged signed block at the start of continuation history. ([27]) Signing and API content checks should be treated as structural acceptance controls. The separate question is whether the transformed state still supports the intended task. Closed-system validation language addresses accuracy, reliability, and consistent intended performance, providing a broader assessment boundary. ([22])

This report proposes three independent verdicts: transport accepted, critical state preserved, and authorized continuation correct. A passing transport verdict must never substitute for the other verdicts. Structure can be checked with schemas, but object fields are optional unless explicitly required. ([28]) Provenance can establish which inputs produced a state without proving that every clinically relevant distinction was retained. ([14])

Constraints that affect implementation

On-demand compaction rejects context_management, stop_sequences, structured-output format, forced tool_choice, and output_config.task_budget.remaining. Its max_tokens budget includes preceding thinking. ([29]) ([30]) Keep these constraints in a versioned request builder, separate from application acceptance criteria. The regression fixture should capture the exact request configuration and expected response handling so parameter changes become visible changes to intended behavior. Shared setup and cleanup are standard fixture capabilities. ([31])

For each acceptance case, record the transformed inputs, required response, observed result, source-state derivation, and schema version. ([3]) ([32]) ([28])

04

Implementation Considerations and Process Changes

Snapshot, continue, and commit

The documented background pattern appends turns while preserving existing history and avoiding another compaction. When the block arrives, only the submitted prefix is replaced; subsequently appended turns stay after it. ([23]) ([33]) The application should represent this as a state machine with full history, compaction pending, candidate ready, validated replacement, and recovery states. These labels are the report's proposed implementation vocabulary.

The following transition discipline is recommended:

  • Snapshot: Persist the exact submitted prefix, its boundary, and the active configuration.
  • Continue: Append new work under the same history lineage while the candidate is pending.
  • Check: Associate the returned candidate with the snapshot that generated it.
  • Validate: Run critical-state probes before allowing consequential continuation.
  • Commit: Publish the replacement and its lineage reference together.
  • Recover: Keep the previous valid branch available if acceptance fails.

Use an application-level version precondition on the commit. A transaction can make several database updates atomic and hide intermediate states from concurrent transactions. It does not supply the application's semantic check or identify a stale compaction candidate automatically. ([25]) Connect the asynchronous request and the foreground work through trace correlation; OpenTelemetry defines identifiers that group spans across processes. ([34]) The durable state record should remain the source of truth, with telemetry assisting diagnosis.

Keep-tail boundaries and thinking

Kept turns must be sent unchanged, including their thinking blocks. Tool calls and results must remain on the same side of the cut. ([24]) A practical qualification case should place an unresolved tool action near the boundary and verify that the application either delays the cut or retains a coherent interaction. Downstream permission checks should still evaluate every action independently. ([15])

For preserved thinking, every intervening compaction must use a preserved-thinking model. Kept messages must immediately follow the summarized messages unchanged, their first role must differ from the preceding role, and the system and non-deferred tools must match across the relevant requests. ([6]) Configuration equality should therefore be an explicit precondition in the proposed state machine. A schema can enforce the presence of configuration identifiers but cannot establish the truth of their contents. ([28])

Use one binding-check model throughout a thinking-enabled compatibility fixture. Retain at least one thinking block; send both compact-2026-09-04 and thinking-binding-controls-2026-08-01; and set thinking.block_binding.prefix_mismatch_behavior to "error". Continue with the signed block, unchanged kept turns, then a new user message. Require HTTP 200 with empty input_transformations only under these preconditions. For a negative control, alter a retained message preceding the thinking block, leaving both the thinking and signed compaction blocks unchanged; require HTTP 400 identifying a conversation-binding mismatch. Reject compatibility if either control fails, then run behavioral probes separately. ([35])

Mid-conversation instructions

Instructions in summarized role: "system" messages stop applying when the block replaces them, even if their content is remembered in the summary. ([36]) Add a protocol-constraint fixture with such a governing instruction. Restate any still-required instruction in a role: "system" message directly after the first new user turn following the kept turns, and retain it in later history. Leave the kept turns unchanged; placing the restatement between the block and kept turns breaks their thinking. ([35]) Verify the expected authorized decision and repeat the positive thinking check. Remembering instruction content is separate from preserving its instruction role.

Evidence attachments and configuration control

Images, documents, container uploads, and fetched URLs inside replaced messages are absent from continuation history. ([37]) This report recommends retaining resolvable evidence references outside the summary and testing rehydration explicitly. W3C provenance provides a vocabulary for relating the derived state to its source entities. ([14]) For clinical-trial migration, the European Medicines Agency (EMA) recommends keeping data, contextual information, and audit trails together. That recommendation informs the archive design; it is not an approval of summary replacement. ([38])

Treat the chosen boundary as a reviewable configuration artifact. Record its source-state relationship and correlate asynchronous activity across processes. When accepting a candidate requires several state updates, commit them atomically so readers cannot observe an incomplete replacement. These are proposed application controls built on documented infrastructure capabilities. ([14]) ([34]) ([25])

F.02
Proposed compaction replacement discipline
01Snapshot

Persist the exact submitted prefix, its boundary, and the active configuration.

02Continue

Append new work under the same history lineage while the candidate is pending.

03Check

Associate the returned candidate with the snapshot that generated it.

04Validate

Run critical-state probes before allowing consequential continuation.

05Commit

Publish the replacement and its lineage reference together.

06Recover

Keep the previous valid branch available if acceptance fails.

Evaluating AI for your business?

Our team helps companies navigate AI strategy, model selection, and implementation.

Get a Free Strategy Call
“

A fluent answer is insufficient evidence of preservation. The table deliberately includes attribution, qualifiers, and executable obligations, because the proposed acceptance boundary is task continuity.

05

Critical-State Inventory for Pharmaceutical Agents

Define preservation before measuring it

The proposed critical-state contract is a typed, versioned inventory of what subsequent decisions require. It should distinguish exact-match identifiers, meaning-sensitive facts, documentary evidence, temporal relationships, and action status. MHRA describes metadata as supplying attributes and context for other data, making it a useful basis for deciding what a prose summary alone cannot represent reliably. ([39]) FDA's guidance also emphasizes copies that preserve record content and meaning. ([40])

Table 2 defines a suggested golden set. Before and after columns describe measurements to collect, not experimental results. Each row must be instantiated with approved fixtures appropriate to the agent's intended use.

T.02
Critical stateBefore compactionAfter compactionProposed release criterion
Case and document identifiersExact authoritative valuesExact values returned by probesNo changed or missing required identifier. ([28])
Adverse-event factsFacts, uncertainty, source associationsMatched facts and qualifiersNo unsupported change to required meaning. ([40])
Citations and evidenceSource identifier, version, locationResolvable matching evidenceNo invented or misattributed supporting source. ([14])
Protocol constraintsApplicable version and conditionsConstraint-aware task decisionSame authorized decision under unchanged facts. ([22])
Temporal updatesOrdered corrections and superseded factsCurrent value plus update historyCorrect latest state without erasing lineage. ([9])
Pending actionsIntent, status, tool operation identifierCorrect next action and statusNo duplicate or unauthorized operation. ([15])
Unknowns and conflictsExplicit unanswered questionsCorrect uncertainty or abstentionNo fabricated resolution. ([9])

A fluent answer is insufficient evidence of preservation. The table deliberately includes attribution, qualifiers, and executable obligations, because the proposed acceptance boundary is task continuity. LongMemEval's evaluation categories include extraction, temporal reasoning, updates, and abstention, supporting this broader test design. ([9]) The International Council for Harmonisation (ICH) recommends reconciliation during data exchange or migration to avoid loss and unintended modifications; this is relevant context for controlled state replacement. ([41])

Separate operational memory from regulated records

Classify each inventory field by authoritative location, intended use, responsible owner, and retention policy. A working summary can be disposable operational memory while its source remains an official record. Where the agent output contributes to a regulated record, the applicable record rules need a documented assessment. ([7]) Applicable closed-system controls include accurate, ready retrieval throughout retention and preservation of prior information when records change. ([42])

For each field, specify:

  • Authority: The source that determines the accepted value.
  • Identity: The immutable identifier and relevant version.
  • Meaning: Qualifiers, units, negation, uncertainty, and relationships.
  • Access: The permission needed to retrieve or act on it.
  • Continuity: The next task that depends on the field.
  • Recovery: The route to reconstruct state when a probe fails.

This inventory should be reviewed as an intended-use specification. OECD's risk-centered lifecycle approach supports scaling validation effort to the significance of the system. ([20]) The contract should explicitly state which fields require exact equality and which need adjudicated semantic equivalence, making a passing result reviewable rather than impressionistic.

For changing fields, include the current value, superseded value, update order, and authoritative source in the fixture; mark required fields in its schema. ([9]) ([14]) ([28])

06

Regression Harness and Context-Retention Testing

F.03
Structural acceptance and task continuity
Structural acceptance
  • Place exactly one unchanged signed block at the start of continuation history.
  • Send kept turns unchanged, including thinking blocks; keep each tool call and result on the same side of the cut.
  • Use a preserved-thinking model for every intervening compaction when preserving thinking.
Task continuity
  • Score both branches against independent authoritative labels to detect errors shared by both.
  • Check identifiers exactly, compare typed fields, and adjudicate paraphrased meaning against its source.
  • Verify the operation's final state as well as the assistant's description of it.

A passing transport verdict must never substitute for critical-state preservation or correct authorized continuation.

Build paired tests against authoritative state

The proposed harness forks an approved transcript into an uncompacted branch and a compacted branch. Both receive identical continuation tasks and tool fixtures. Score each branch against independent authoritative labels; comparing branches alone could accept an error shared by both. The National Institute of Standards and Technology (NIST) Generative AI Profile recommends ground-truth comparison and empirical assessment of capability claims. ([43]) Test cases should specify input and expected behavior rather than merely confirm that a function returned a response. ([32])

Qualification should include these test families:

  • Fact recall: Probe every required identifier and meaning-sensitive field.
  • Citation preservation: Check source identity, version, location, and supporting passage.
  • Decision continuity: Compare the next task outcome against the approved answer.
  • Position variation: Move critical evidence among early, middle, and recent turns. ([8])
  • Update handling: Insert later corrections and test which value governs. ([9])
  • Repeated replacement: Evaluate several successive summaries, not only the first. ([2])
  • Tool sequencing: Exercise action status and permission checks. ([15])
  • Boundary variation: Repeat cases with different retained-tail cuts and history lengths. ([31])

Treat the uncompressed branch as a baseline, not an oracle. Lost in the Middle shows that relevant-information position can change performance, including in long-context models studied in that work. ([8]) A compacted branch may improve one retrieval task while degrading another. Report results by risk category so gains in ordinary recall cannot mask changed case identity or lost constraints.

Make the oracle explicit

Use exact equality for identifiers, typed comparisons for structured fields, and source-grounded adjudication for paraphrased meaning. JavaScript Object Notation (JSON) Schema should mark required properties explicitly and decide whether extra properties are allowed. ([28]) ([44]) Check each asserted onset date, causal qualifier, and source relationship against the independent expected answer.

Partition tests by indexing, retrieval, and reading, the memory-system stages described in the LongMemEval paper. ([45]) If an archived source cannot be found, that is a retrieval problem; if it is retrieved but ignored, that is a reading problem. Use this separation to avoid blaming summary quality for every downstream deviation. An evidence lineage record links the relevant source entity to the derived state and observed answer. ([3])

A corrected-case fixture (Hypothetical Example)

Construct a synthetic safety case with an initial reported fact, a later correction, an unresolved follow-up, and a draft action that has not been executed. Before compaction, ask which fact is current and which action is pending. After compaction, repeat those probes and request the next workflow step. The proposed expected result preserves the correction, cites its source, and keeps the action unexecuted until authorized. Update handling and abstention are independently recognized evaluation dimensions; downstream authorization remains a separate control. ([9]) ([15])

Vary the correction's position and the cut boundary while reusing the same fixture. Shared preparation and cleanup make this practical, but the test must retain an independent expected answer. ([31]) No preservation percentage is implied by this example. Release thresholds should be approved through a documented risk assessment and should explicitly identify which failures prevent deployment. ([10])

Include retrieval tasks that require combining evidence, rather than testing isolated strings exclusively. RULER includes multi-hop and aggregation tasks, while the staged memory model distinguishes finding evidence from reading it. Expected outcomes should still be expressed as concrete test assertions for the actual agent. ([46]) ([45]) ([32])

07

Tool Use, Failure Modes, and Recovery

Preserve action semantics across replacement

The proposed tool-use suite distinguishes an intended call, an issued call, a completed result, and a committed business action. Each transition should have durable state outside the summary. Authorization belongs in downstream systems rather than in the model's recollection of a permission grant. ([15]) Applicable closed-system controls also include authority checks. ([47])

Use separate fixtures for a delayed result, a retry after an uncertain response, a cancelled action, and a resumed session. Trace identifiers can connect the foreground task to the asynchronous summarization activity. ([34]) The oracle should verify the operation's final state as well as the assistant's description of it. A correct narrative cannot substitute for a database or destination-system check.

Table 3 separates documented API outcomes from proposed application responses. The semantic test row is an application failure condition, not an Anthropic error code.

T.03
Observed conditionMeaningProposed response
400 compaction_block_misplaced, compaction_signature_invalid, or compaction_content_mismatchPlacement or block validation failed. ([48])Repair serialization; retain the previous valid history.
529 compaction_unavailableTransient summarization problem. ([49])Bounded retry with backoff; continue only within safe operating limits.
200 without a usable blockSummarization did not necessarily produce a candidate. ([50])Inspect stop_reason; reject unusable candidates, raise the budget for max_tokens, or shorten submitted input for model_context_window_exceeded.
Continuation thinking mismatchCompatibility may fail or drop thinking. ([51])Reject the branch until the continuation policy passes.
Critical-state probe failsApplication acceptance criterion failed. ([32])Quarantine candidate, restore the valid branch, review the deviation.

Retry policy should follow the failure class. Repeating an invalid block does not change its validity, while a transient availability response may justify a bounded retry. Recovery should preserve enough context to explain which branch was accepted and why. Provenance records can relate the candidate and retained branch to the same source history. ([14]) Database rollback is a useful local implementation mechanism, although it does not reverse an external business action. ([52])

Test recovery as deliberately as success

The recovery plan should contain:

  • Retry limits: Bound attempts and elapsed time according to workflow needs.
  • Candidate isolation: Prevent an unaccepted summary from becoming active state.
  • Branch restoration: Reload the last validated history and its configuration.
  • Operation reconciliation: Check downstream action status before reissuing work.
  • Escalation: Route unresolved consequential decisions to the designated reviewer.

These are proposed operational controls. OWASP recommends approval for high-impact actions, and ICH recommends contingency procedures for essential data. ([53]) ([41]) Exercise the recovery path with deterministic fixtures rather than waiting for production exceptions. Test preparation should include the state needed to reproduce each failure and cleanup that restores the environment. ([31])

Recovery evidence should identify the interrupted transformation, the retained branch, and the operation's authorized status. Database transactions support local rollback, but model-side narration must not determine access. Keep the source-state relationship available to the reviewer when deciding how to resume. ([52]) ([15]) ([3])

08

Audit Evidence and Operational Ownership

Record the transformation and the decision

The proposed compaction audit packet should make the replacement reconstructable. W3C PROV models provenance around data entities, activities, and responsible participants, while its data model defines derivation as transformation between entities. ([14]) ([3]) This offers a useful organizing scheme without establishing regulated electronic-signature compliance.

For each accepted or rejected candidate, capture:

  • Source lineage: Transcript snapshot identity, boundary, and authoritative references.
  • Configuration: Model, platform, request builder, system, tools, and test-contract versions.
  • Candidate evidence: Returned block, response status, request identifier, and usage record.
  • Compatibility verdict: Serialization checks and continuation-thinking observations.
  • Semantic verdict: Probe outputs, expected labels, deviations, and adjudication.
  • Operational outcome: Accepted branch, reviewer decision, recovery action, and downstream state.

Store the block without editing its contents, but keep the explanatory packet separately. The application can link related spans across traces using OpenTelemetry links. ([54]) Durable evidence should retain the relationship between inputs, transformation, and continuation; traces alone are not a substitute for the record archive. WHO guidance recommends retaining electronic records with associated metadata, including audit trails and electronic signatures. ([55])

Establish scope, access, and review ownership

The validation owner should approve intended use and release criteria; the application owner should control request changes and recovery; the workflow owner should adjudicate domain semantics. These are suggested responsibilities. OECD's lifecycle framing supports treating validation and operation as connected activities rather than a one-time integration check. ([21]) NIST's AI Risk Management Framework is voluntary, so using its terminology should not be presented as a certification. ([56])

Where closed-system requirements apply, prior recorded information must not be obscured by changes, and audit documentation must be retained with the associated records for the required period. ([42]) EMA's guidance similarly recommends preserving contextual information and audit trails during migration. ([38]) Decide whether the working summary belongs in the regulated record set and document the rationale, rather than assigning a universal retention duration to every conversation.

Access to the archive should preserve the same practical permission boundary as access to the source. Downstream authorization guidance supports enforcing this outside the language model. ([15]) Avoid logging sensitive content indiscriminately; choose the minimum evidence needed for reproducibility and review under the workflow's approved information controls. Provenance relationships can refer to controlled entities without requiring every telemetry event to duplicate their contents. ([14])

Assign each evidence item to the durable record packet or operational telemetry, and record its source reference, trace link, and retention owner. ([14]) ([54]) ([11])

09

Data Analysis and Evidence

What existing research measures

The relevant public evidence concerns evaluation design rather than validated pharmaceutical compaction performance. RULER evaluates 17 long-context models using 13 representative tasks, extending retrieval with multi-hop tracing and aggregation. ([46]) LongMemEval contains 500 curated questions and evaluates extraction, multi-session reasoning, temporal reasoning, updates, and abstention. ([57]) Lost in the Middle studies multi-document question answering and key-value retrieval, with position-dependent performance. ([8])

The August 2, 2026 version of Context Compaction Theory examines a single compaction invocation and distinguishes retained subsets from bounded generated summaries. ([2]) These studies should not be pooled into a claimed success rate for Anthropic's on-demand feature. Their tasks, models, and methods differ. Their practical contribution here is a test taxonomy: probe retrieval, compositional reasoning, changing facts, uncertainty, and successive state transformations separately. ([46]) ([45])

Measure retention and task continuity

The proposed report card should calculate required-fact recall as correctly recovered required facts divided by applicable required facts. Calculate citation preservation as correctly associated, resolvable evidence references divided by required references. Measure decision accuracy, pending-action continuity, and unsupported additions separately. These are analyst-defined measures, not published results. Independent expected labels and concrete assertions make their denominators auditable. ([32]) ([28])

Report before and after values for every golden-set row, with the number of applicable cases, excluded cases, and adjudicated disagreements. Split the results by workflow, risk category, history position, tail boundary, and compaction count. Position-sensitive tests are motivated by the long-context literature; indexing, retrieval, and reading should be diagnosed independently. ([8]) ([45]) State the conditions under which the evidence applies instead of reporting one blended score as a universal preservation guarantee.

Calculate economics from actual API usage

For an uncached continuation, the proposed input-only calculation is: uncompacted input tokens minus compacted input tokens, multiplied by the applicable input price per token. Sonnet 5.5's current input rate is $2 per million tokens. ([12]) This calculation is a reporting formula, not a measured saving. It excludes summarization and output costs and should be used only with actual usage from comparable continuations.

On-demand generation is billed; its cost is in usage.iterations, while top-level input and output tokens are zero. ([58]) Cached input totals combine input_tokens, cache_creation_input_tokens, and cache_read_input_tokens. ([59]) Sonnet 5.5 cache rates are $2.50 for five-minute writes, $4 for one-hour writes, and $0.20 for reads per million tokens. ([12]) Account for each category at its own rate, plus output, retries, and additional retrieval work.

Token counting is an estimate, is free within rate limits, and does not simulate cache behavior. ([13]) Therefore a smaller estimated prompt is not a measured billing result. The economic experiment should compare complete sessions under the same task and cache policy, then relate cost to the already approved semantic outcomes. OECD's risk-based scaling principle supports prioritizing acceptable performance over an isolated efficiency figure. ([20])

A benchmark description is useful only within its method. RULER's synthetic tasks broaden retrieval testing; LongMemEval's curated conversational questions cover changing and temporal information. They motivate different fixture families, not a conversion factor from context length to pharmaceutical correctness. This report deliberately keeps those populations separate from the proposed workflow tests. ([46]) ([57])

To interpret a failed probe, retain the expected answer, relevant source, actual output, and failure classification. A specific input-response assertion and an explicit state schema support reproducibility; the indexing, retrieval, and reading distinction supports diagnosis. This proposed report card should accompany cost measurements rather than be replaced by them. ([32]) ([28]) ([45])

10

Implications and Future Directions

The recommended adoption decision is conditional: proceed only when the state contract, paired regression, and recovery evidence support the specific workflow. The following decision sequence is proposed, rather than presented as a regulatory checklist:

  • Scope the work: Establish whether agent state or outputs contribute to required records. ([7])
  • Check compatibility: Qualify the selected platform, model, and request builder.
  • Define critical state: Approve authoritative labels and semantic acceptance rules. ([22])
  • Test the transformation: Run paired recall, citation, decision, and boundary cases. ([32])
  • Test continuation: Verify authorization and downstream operation state. ([15])
  • Prove recovery: Reconstruct the valid branch and reconcile pending work. ([25])
  • Review change: Requalify affected cases when configuration or intended use changes. ([21])

For a scientific reading assistant, the contract might emphasize source-grounded retrieval; for a workflow agent, operation status and permission can be decisive. Those are intended-use choices, not measured product distinctions. LongMemEval's staged memory-system analysis provides a useful way to separate source discovery from interpretation when diagnosing these different workloads. ([45])

Future vendor changes should be incorporated through a controlled delta review. Preserve the earlier evidence, identify altered interfaces or behavior, and select affected fixtures. The theory paper's single-invocation scope also leaves repeated compaction as a separate empirical question for the application. ([2]) No inferred guarantee should bridge that gap.

IntuitionLabs' stated measurement perspective includes quality, risk signals, reliability, and support burden before expansion. ([60]) Applied here, that means scaling only after both task performance and operational ownership are demonstrated. NIST's voluntary risk framework offers broader governance context; it does not replace workflow-specific qualification. ([56])

Document the deployment organization's record scope, acceptance thresholds, reviewer authority, and lifecycle review owner. ([56]) ([21])

11

Conclusion

Adopt Claude API on-demand compaction when the qualification package supports the intended workflow: an approved critical-state inventory, independent expected answers, paired before-and-after tests, passing continuation controls, and exercised recovery. Record which state was submitted, which continuation was tested, which critical facts survived, who reviewed the result, and how exceptions were resolved. ([3]) ([32]) ([15]) ([14])

Assess the workflow's record scope and risk before release. Where applicable, verify accurate copies, recoverable history, and reviewable changes. ([11]) Base the economic decision on actual session usage and category-specific costs, reported alongside the test outcomes.

The publisher

About IntuitionLabs

Build practical AI for pharma and biotech with IntuitionLabs. We help life-science teams turn complex information and workflows into useful software, governed knowledge systems and AI tools.

IntuitionLabs is an AI consulting, custom software development and data engineering firm serving pharmaceutical, biotechnology, medical-device and other life-science organizations. We work with clinical, regulatory, medical-affairs, commercial, quality and IT teams to connect technology decisions with the work people need to accomplish.

AI consulting and adoption

Our AI enablement services cover readiness assessments, use-case selection, governance and policies, team workshops, adoption measurement and ongoing advisory support. We help organizations structure the information layer behind AI: source material, context, permissions and maintained knowledge that make generated answers useful and reviewable. Private LLM inference and hosted AI options support teams evaluating how to operate AI with appropriate control over their data and infrastructure.

Software, data and life-science workflows

IntuitionLabs develops custom software for pharma and biotech, integrates enterprise systems, and builds data engineering and business intelligence solutions. Areas of focus include AI agents, regulatory research, medical writing, medical affairs, CMC information, competitive intelligence and clinical-document workflows. Our eTMF intelligence work includes cross-system reconciliation and inspection-readiness support.

Enterprise platforms and regulated delivery

We provide Veeva services, application support, managed services, integrations and custom applications, alongside enterprise content work involving platforms such as Egnyte. For regulated workflows, our services include GxP enablement, computer-system validation and software development addressing 21 CFR Part 11 requirements. The applicable controls, validation responsibilities and acceptance criteria are defined for each engagement.

Work with IntuitionLabs

Explore AI enablement, pharma and biotech software development, data engineering and BI, and Veeva services. Contact IntuitionLabs to discuss your workflow, information sources and implementation needs.

IntuitionLabs publishes educational research to help life-science teams make informed technology decisions. Coverage of a product or organization does not imply a client relationship, endorsement or partnership.

Sources / 60

Get a Free AI Cost Estimate

Tell us about your use case and we'll provide a personalized cost analysis.

Ready to implement AI at scale?

From proof-of-concept to production, we help enterprises deploy AI solutions that deliver measurable ROI.

Book a Free Consultation

Turn This Insight into a Working Life-Sciences Workflow

IntuitionLabs connects governed information, specialist implementation, role-based adoption, and measured value.

Disclaimer

The information contained in this document is provided for educational and informational purposes only. We make no representations or warranties of any kind, express or implied, about the completeness, accuracy, reliability, suitability, or availability of the information contained herein. Any reliance you place on such information is strictly at your own risk. In no event will IntuitionLabs.ai or its representatives be liable for any loss or damage including without limitation, indirect or consequential loss or damage, or any loss or damage whatsoever arising from the use of information presented in this document. This document may contain content generated with the assistance of artificial intelligence technologies. AI-generated content may contain errors, omissions, or inaccuracies. Readers are advised to independently verify any critical information before acting upon it. All product names, logos, brands, trademarks, and registered trademarks mentioned in this document are the property of their respective owners. All company, product, and service names used in this document are for identification purposes only. Use of these names, logos, trademarks, and brands does not imply endorsement by the respective trademark holders. IntuitionLabs.ai is an AI software development company specializing in helping life-science companies implement and leverage artificial intelligence solutions. Founded in 2023 by Adrien Laurent and based in San Jose, California. This document does not constitute professional or legal advice. For specific guidance related to your business needs, please consult with appropriate qualified professionals.

Related Articles

Need help with AI?

© 2026 IntuitionLabs. All rights reserved.