FlexHarness is a modular toolbox for model inference, agents, managed sessions, chat, OCR, KVM control and provider capabilities. It brings SmartAI, SmartAgent, SmartChat, SmartKVM and SmartOCR's AI helpers into one repository through four focused packages published with tspublish.
Issue Reporting and Security
For reporting bugs, issues, or security vulnerabilities, please visit community.foss.global/. This is the central community hub for all issue reporting. Developers who sign and comply with our contribution agreement and go through identification can also get a code.foss.global/ account to submit Pull Requests directly.
Install and choose an entrypoint
The toolbox publishes four packages under @modelprofile.com, with one shared
release version. Integrations use subpath imports instead of additional package
names.
| Package | Use it for |
|---|---|
flexharness |
Managed sessions, permissions, tools, chat, OCR, documents and KVM |
flexharness-agent |
Standalone runAgent and AgentSession; portable /runner |
flexharness-models |
Model contracts, ModelRegistry, AI SDK helpers and caching |
flexharness-providers |
All eight model providers, authentication, accounts, media and research |
For a managed agent with model providers:
pnpm add @modelprofile.com/flexharness @modelprofile.com/flexharness-providers
For inference without an agent:
pnpm add @modelprofile.com/flexharness-models @modelprofile.com/flexharness-providers
import { ModelRegistry, generateText } from '@modelprofile.com/flexharness-models';
import { createOpenAiModelProvider } from '@modelprofile.com/flexharness-providers/openai';
const models = new ModelRegistry().register(createOpenAiModelProvider());
const setup = models.getModelSetup({
provider: 'openai', model: 'gpt-5.5', apiKey: process.env.OPENAI_API_KEY,
});
const result = await generateText({ ...setup, prompt: 'Hello' });
For a standalone agent, install @modelprofile.com/flexharness-agent and pass the
same setup to runAgent({ ...setup, prompt: 'Hello' }). Node.js 24 or newer is
required for Node entrypoints. Each application owns its model registry; imports
do not register providers globally.
Harness subpaths
All entries below belong to @modelprofile.com/flexharness. The root export stays
focused on managed sessions. Importing it does not load UI, PDF, browser automation
or provider SDKs.
| Import suffix | Capability | Additional install |
|---|---|---|
/tools |
Host-supplied execution contexts, filesystem (with opt-in edit_file), shell, HTTP and JSON tools |
None |
/tools/node |
Local filesystem and process tools, file-backed job adapters, InMemoryWorkspaceReversion for turn undo of local files |
None |
/browser |
Run-scoped SmartBrowser framed-client adapter for Flex tool providers | @push.rocks/smartbrowser |
/compaction |
Model-driven conversation compaction | None |
/media |
Vision recipes using an injected model | None |
/ocr |
ImageOcr, OCR contracts and Mistral transport |
None |
/chat |
Portable streaming ChatSession, history and usage |
None |
/chat/cli |
Ink/React terminal chat | ink ink-text-input react; TypeScript: @types/react |
/chat/web |
Lit flexchat-window, flexchat-message, flexchat-input |
lit |
/chat/harness |
FlexHarnessTranscript, intents, the chat wire contract and FlexHarnessChatClient for the @design.estate/dees-catalog harness components, browser-safe |
@design.estate/dees-catalog ^17.0.0 || ^19.5.0 || ^20.0.0 || ^21.0.0 (types only) |
/chat/harness/server |
FlexHarnessChatHub: conversations, steering, permissions, undo and tools that run in the tab, for browser tabs over any transport |
@design.estate/dees-catalog (types only) |
/chat/harness/typedsocket |
The hub and the chat client over TypedSocket | @api.global/typedsocket ^9.1.0, @api.global/typedrequest ^9.1.0, @api.global/typedrequest-interfaces ^7.1.0 |
/kvm |
Browser control, terminal framing and OCR observation | puppeteer |
/documents |
PDF and image documents, extractTextFromPdf |
@push.rocks/smartpdf; for image documents sharp |
/mcp |
MCP clients, AI SDK tool conversion and createMcpToolCatalog, the tool catalog of MCP tools |
@modelcontextprotocol/sdk |
/migration |
Existing versioned harness migrations | None |
/stores/nosqldb |
NoSqlFlexHarnessStores: cross-process CAS stores on @lossless.org/client/nosqldb |
@lossless.org/client |
/stores/testing |
createFlexStoreContractTests: the contract every built-in IFlexHarnessStores passes, to run against a custom store with any test runner |
None |
/typedrequest |
describeTypedRequests, catalogFromTypedRequests: the tool catalog of a typedrequest API from its TypeScript sources, answered by its router in process |
@api.global/typedopenapi to describe an API |
The integrations in the last column use optional peer dependencies. Install them only for the entrypoints your project imports, for example:
pnpm add @modelprofile.com/flexharness lit
For a managed SmartBrowser resource, install its optional browser peer:
pnpm add @modelprofile.com/flexharness @push.rocks/smartbrowser
import { ChatSession } from '@modelprofile.com/flexharness/chat';
import '@modelprofile.com/flexharness/chat/web';
Chat uses the portable flexharness-agent/runner. A browser application can provide
a model backed by its own server transport, or supply run in
IChatSessionOptions. The web element accepts an IChatSession and maintains its
own subscription. Full Node AgentSession remains at the agent package root.
Durable session token metrics
Read await harness.getSessionMetrics(scopeId, sessionId) for bounded, durable
accounting independent of transcript visibility. It reads fixed-size projection
state rather than summing assistant messages or scanning event/archive/child history.
A warm read joins outstanding projection writes and surfaces accounting failures.
| Field | Meaning |
|---|---|
reportedLifetimeUsedTokens |
Sum of durably settled per-call reports. Incomplete accounting makes this a lower bound. |
reportedLifetimeUsedTokensExact |
Whether the sum of those reports is exact; invalid reports and safe-integer overflow make it inexact. |
lifetimeUsageComplete |
Every owned call is accounted for and no accounting operation remains in flight. |
lifetimeUsedTokens |
Present only when the lifetime total is complete. Zero is a valid total. |
inFlight |
A generation or actual compactor invocation has not settled its accounting yet. |
lastObservedInputTokens, lastObservedInputTokensAt |
Input count and observation time of the last settled generation call with complete input/output totals. |
lastObservedInputTokensFreshness |
Always historical. This is not current context size: output, steering, compaction and reversion can change context. |
reportedLifetimeInputTokens |
Sum of the settled per-call reports' input tokens, prompt-cache reads included. |
reportedLifetimeCacheReadTokens, reportedLifetimeCacheWriteTokens |
Sums of the settled per-call reports' input tokens read from and written to the prompt cache. |
lifetimeTokenBreakdownComplete |
The three sums cover every call lifetimeUsedTokens covers: lifetime usage is complete and the session's calls were counted by input and cache since it began. Otherwise they are lower bounds. |
Accounting covers this session's generation calls and reported compaction calls,
including failure, cancellation and retries. Child sessions own their own spend;
tool-owned model calls are not included. Do not sum cumulative assistant-message
usage to reconstruct lifetime spend. Missing/partial reports remain incomplete;
unsafe totals saturate at Number.MAX_SAFE_INTEGER without claiming exactness.
A compactor that reports no calls cannot prove zero spend.
New sessions start complete at zero. Existing/imported sessions with unknown prior spend start incomplete. Write-ahead operation markers make interrupted recovery conservative: unresolved operations are cleared without certifying their unknown spend. Undo, redo, compaction and transcript archiving never subtract lifetime spend. Importing visible history does not reconstruct historical billing.
Sessions whose metrics were written before input and cache tokens were counted count them from their next settled call on and never report the breakdown complete; their lifetime total keeps its completeness.
Conversational runs that exhaust maxSteps with tool work or unseen input remaining
fail with FlexHarnessStepLimitError (FLEX_STEP_LIMIT), rather than being marked
completed with an empty answer. Standalone Agent callers retain tool-only results
and can inspect terminationReason: 'model' | 'step-limit'.
Provider subpaths
Install @modelprofile.com/flexharness-providers once. Its root exports all provider
factories; use a provider subpath to import just that implementation.
| Import suffix | Capability |
|---|---|
/anthropic, /openai, /google, /groq, /mistral, /xai, /perplexity, /ollama |
Model adapters and their request types |
/openai (ChatGPT) |
Inference and the model catalog with access resolved by an account authority |
/openai (web search) |
createOpenAiWebSearchTool(): OpenAI's provider-executed web search as a run tool |
/openai (text files) |
Text file parts (text/*, JSON, XML, YAML, any type with a charset) are sent as text with their file name, on both APIs and all connections |
/auth |
ChatGPT browser/device authentication, refresh and model connections |
/accounts |
Account lifecycle, model catalogs, rate limits and credential envelopes |
/auth/files |
Explicit interoperability with external credential files |
/media |
OpenAI audio and image generation |
/research |
Anthropic research utilities |
/anthropic-messages, /anthropic-messages/node |
Serves the Anthropic Messages API over FlexHarness models, for Anthropic clients such as Claude Code |
Provider SDKs are installed together. Authentication, file access and vendor media
are separate entrypoints and are not imported by the provider factory root. For
ChatGPT authentication, pass
connection: createOpenAiChatGptModelConnection(credentials) from /auth to the
OpenAI model options. API-key inference needs no authentication setup.
HTTP gateways forwarding already serialized Responses JSON can call
normalizeOpenAiChatGptResponsesRequest(body) from
@modelprofile.com/flexharness-providers/openai before sending it through a
resolving connection's settings.fetch. It lifts leading system/developer text
into instructions when no instruction prefix was supplied, preserves
later instruction order as developer messages, and removes unsupported
prompt_cache_retention and max_output_tokens. Explicit instructions and
other request fields are preserved. Non-text instruction prefixes are rejected;
the caller's body is never mutated. Authentication, account selection, storage and streaming policy remain
with their existing owners.
Usage and rate limits
The OpenAI and Anthropic adapters turn a provider's limit response into one typed error,
ModelLimitError from @modelprofile.com/flexharness-models, and keep the provider failure as
its cause. Its limit is a bounded IModelLimitInfo:
| Field | Meaning |
|---|---|
kind |
usage_limit: the account's quota is spent until it resets, so no retry succeeds earlier. rate_limit: a short throttle that a retry after the advertised delay may pass. |
provider |
openai or anthropic. |
resetsAt |
When the limit resets, in epoch milliseconds. |
retryAfterMs |
The delay the provider asked for (retry-after-ms, retry-after). |
planType, window |
The account plan and the exhausted window, for example plus and 5h. |
Each adapter reads its own dialect. The ChatGPT connections classify Codex
usage_limit_reached and usage_not_included (HTTP 429 bodies and failed streams), taking the
reset from resets_at, resets_in_seconds or the exhausted x-codex-* window. The OpenAI API
dialect treats insufficient_quota as a usage limit and any other 429 as a rate limit. Anthropic
treats anthropic-ratelimit-unified-status: rejected as a usage limit until the unified reset,
naming the 5h or 7d window whose reset it matches, and a 429 rate_limit_error as a rate
limit. A limit a stream reports is the same ModelLimitError, whether it arrives before the answer
begins or after text has streamed. The AI SDK keeps only the
code and message of an OpenAI stream error frame, so a ChatGPT limit reported by a stream takes
its reset and window from the stream response's x-codex-* headers, and a stream carries no
retry-after. isModelLimitError() recognizes the error across package copies. Because the typed error
is not an APICallError, the AI SDK does not retry it; the agent runtime owns retries.
describeModelFailure(error) names why a model call failed without provider response text,
headers or credentials: { kind, status?, providerCode?, limit? }, where kind is limit, http
(an HTTP error answer and its status), stream (an error a successful response stream reported
before or after output began, with the machine-readable providerCode the provider states and no
status, since the answer itself succeeded), response (an
answer that could not be processed), network, aborted (a timeout signal included) or unknown.
It follows cause and the last attempt of a retried call, so a host can say "HTTP 503" or "usage
limit" without depending on the AI SDK's error classes.
An OpenAI connection may add its own dialect with classifyLimit; it is consulted before the
OpenAI API dialect. Other adapters can adopt the same contract with
createModelLimitMiddleware({ provider, classify }), whose classifier receives the failure's HTTP
status, lower-case headers, parsed error body or the provider's stream error frame with the status
the provider states for it, and observedAt, the time
relative reset delays count from. A delay or reset time outside the range a Date can represent
is dropped from the limit.
Migration
Replace the withdrawn 6.x component imports with the subpaths above. Provider
imports change from individual provider packages to
@modelprofile.com/flexharness-providers/<provider>. Keep direct agent and models
imports and upgrade them to the same release version. The removed package names
are not republished.
| Previous API | Replacement |
|---|---|
| SmartAI models, caching and AI SDK contracts | flexharness-models and explicit ModelRegistry registrations |
| SmartAI provider factories | flexharness-providers/<provider> |
| SmartAI authentication, account adapters and external credential files | flexharness-providers/auth, /accounts, /auth/files |
| SmartAI vision, document and OCR helpers | flexharness/media, /documents, /ocr |
| SmartAI audio, image and research helpers | flexharness-providers/media, /research |
| SmartAgent runtime, events and persistence contracts | flexharness-agent |
| SmartAgent tools, local tools, compaction and MCP | flexharness/tools, /tools/node, /compaction, /mcp |
| SmartChat sessions, CLI and web components | flexharness/chat, /chat/cli, /chat/web |
| SmartKVM | flexharness/kvm: BrowserKvm, KvmTerminal, createKvmTools |
| SmartOCR image AI helper | ImageOcr.recognizeImageBytes from flexharness/ocr |
| SmartOCR PDF AI helper | extractTextFromPdf from flexharness/documents |
For image OCR, rename smartAiOcrEngine to engine and mistralOcrOptions to
mistralOptions; pass the API key explicitly. Native searchable-PDF processing
through SmartOcr.processPdfBuffer remains in SmartOCR.
Existing ISmartAi* and TSmartAi* contracts, durable event/job schemas, harness
projections, credential envelopes and external credential paths remain unchanged.
The OpenAI account SmartAiProviderRegistry is separate from inference's
ModelRegistry. The package consolidation requires no data migration.
Developing and releasing
Four named tspublish.json descriptors own the published packages. Other folder
descriptors specify compilation order only. Each package explicitly declares its
owned folders, exports and required or optional peer dependencies. Source and
compiled declaration mappings support development without workspace links.
Run pnpm build, pnpm test and pnpm run check:test. Tests install packed
artifacts in disposable directories to check minimal dependency graphs, every
public subpath and browser imports. Live provider and browser authentication tests
remain opt-in in test_integration/.
The configured GitZone release prepares four packages, records their exact artifacts, and publishes models before its dependents to npmjs and Verdaccio. Third-party OpenAI Codex notices accompany the providers package.
Core Setup
import {
FlexHarness,
JsonFileFlexHarnessStores,
type IFlexResolvedModel,
type TFlexAgentToolSet,
} from '@modelprofile.com/flexharness';
interface IProjectScope {
projectRoot: string;
}
const stores = new JsonFileFlexHarnessStores({
directory: '/var/lib/my-app/model-sessions',
});
const harness = new FlexHarness<IProjectScope>({
scopeResolver: {
async resolveScope(scopeId) {
const project = await projectRegistry.get(scopeId);
return {
// Aliases that resolve to this same key share sessions and save ordering.
storageKey: project.accountAndProjectKey,
scope: { projectRoot: project.root },
};
},
},
modelResolver: {
async resolveModel({ scope, modelHint, signal }): Promise<IFlexResolvedModel> {
const configured = await modelRegistry.resolve({ scope, modelHint, signal });
return {
model: configured.model,
identity: {
provider: configured.providerId,
model: configured.modelId,
displayName: configured.label,
},
providerOptions: configured.providerOptions,
};
},
},
toolProvider: {
async provideTools(context) {
const tools: TFlexAgentToolSet = await createProjectTools({
root: context.scope.projectRoot,
signal: context.signal,
requestPermission: context.requestPermission,
});
return {
tools,
close: async () => closeProjectTools(tools),
};
},
},
stores,
builtInTools: {
renameSession: true,
projectManagement: {
// task, goal, and scratchpad default to true when this block exists.
},
},
toolOutputLimits: {
maxDepth: 12,
maxBytes: 256 * 1024,
},
callbackLimits: {
maxEvents: 10_000,
maxOutputBytes: 1024 * 1024,
maxParts: 2_000,
},
promptQueueLimits: {
maxOutstandingPromptsPerSession: 16,
maxOutstandingBytesPerSession: 64 * 1024 * 1024,
maxPendingAdmissions: 64,
maxPendingAdmissionBytes: 128 * 1024 * 1024,
maxTerminalEntriesPerSession: 64,
},
reversionLimits: {
maxCompletedTurns: 100,
maxSegments: 300,
maxExcludedRunIds: 1000,
maxPendingReversionReleases: 1000,
},
sessionImportLimits: {
maxOpenImports: 2,
maxMessagesPerSession: 2048,
maxSessionBytes: 64 * 1024 * 1024,
maxPageBytes: 1024 * 1024,
stagedImportIdleTimeoutMs: 5 * 60 * 1000,
},
resultLimits: {
maxStoredBytes: 8 * 1024 * 1024,
maxResultsPerSession: 32,
},
subagents: [
{
name: 'researcher',
description: 'Research a focused question and return one final answer.',
modelHint: 'reasoning-model',
system: 'Investigate the assigned question. Return a concise evidence-based answer.',
maxSteps: 8,
},
],
maxSubagentDepth: 6,
maxSubagentCallsPerRun: 32,
externalErrorProjector: (_error, context) => ({
name: 'ModelOperationError',
message: `The ${context.source} operation failed.`,
code: 'MODEL_OPERATION_FAILED',
}),
});
modelRegistry, projectRegistry, createProjectTools, and closeProjectTools in this example are application-owned integrations. FlexHarness passes the same run AbortSignal to the model resolver and tool provider. Every model-resolver, application tool-provider, and resource tool-provider context also carries the required canonical sessionGenerationId and sessionGenerationSequence, allowing host operations to authorize the exact session generation rather than a reusable session ID alone.
Resource Tool Providers
resourceToolProviderResolver composes zero or more resource-owned providers with the existing application toolProvider. The resolver runs fresh for every prompt and returns the current resource attachment descriptors:
resourceToolProviderResolver: {
async resolveResourceToolProviders({
scope,
sessionId,
sessionGenerationId,
sessionGenerationSequence,
runId,
signal,
}) {
const attachments = await resourceRegistry.listAttached({
scope,
sessionId,
sessionGenerationId,
sessionGenerationSequence,
runId,
signal,
});
return attachments.map((attachment) => ({
resourceId: attachment.resourceId,
attachmentRevision: attachment.attachmentRevision,
provider: createResourceToolProvider(attachment),
}));
},
},
Each descriptor uses the existing IFlexToolProvider<TScope> contract. Its provider receives the normal run context and must return a fresh run-scoped handle. The resource resolver context carries the same required canonical session generation as the tool-provider contexts. The original toolProvider remains optional and its tool names remain unchanged. Resource tool names are deterministic and bounded:
resourceIdentityis the lowercase hexadecimal SHA-256 ofJSON.stringify([resourceId, attachmentRevision]).- The namespace is
resource_plus the first 16 digest characters. - The exposed name is
<namespace>__<stem>__<toolDigest>.stemreplaces characters outside[A-Za-z0-9_-]with_, keeps the first 16 characters, and falls back totool;toolDigestis the first 12 lowercase hexadecimal characters of SHA-256 over the original tool name.
The resolver accepts at most 128 descriptors per run. resourceId must be non-empty and at most 512 UTF-8 bytes, attachmentRevision must be a non-negative safe integer, and each original resource tool name must be non-empty and at most 512 UTF-8 bytes. FlexHarness rejects duplicate resourceId values even across revisions, duplicate derived namespaces, and duplicate final exposed tool names before model execution. Descriptor identity and namespace validation completes before any application or resource provider is acquired.
Resource permission requests are scoped with the complete 64-character resourceIdentity, not the shortened tool namespace. FlexHarness rewrites kind to resource.<resourceIdentity>.<providerKind> and an optional rememberKey to resource:<resourceIdentity>:<providerRememberKey>. Harness-owned metadata contains resourceId, attachmentRevision, resourceIdentity, and toolNamespace; provider metadata is nested under providerMetadata, so it cannot override attachment identity.
FlexHarness owns every acquired handle. Normal close and partial-failure cleanup run in reverse acquisition order, attempt every handle, aggregate multiple failures, and retain failed cleanup for retirement or disposal retry. Cancellation uses the same path. If model resolution fails while a resource provider is still settling, a late returned handle remains tracked and disposal waits for its closure. Application and resource providers may not define a harness built-in name while that built-in is enabled for the current run. Disabled names are not reserved.
Tool Catalog
Every tool in tools sends its description and JSON Schema to the model on every step. A large API surface belongs in the handle's catalog instead: the model sees three fixed tools and reads the API as a typed SDK, only for the methods it needs.
| Tool | Input | Returns |
|---|---|---|
search_tools |
query, namespace, limit, offset |
Ranked path - summary lines and the declaration of a clear top match, or a namespace listing like ls |
describe_tools |
methods (dotted paths), types |
TypeScript declarations of the methods and the named types they use |
execute |
method (dotted path), params |
The method's result |
run_code (code mode) |
code, a TypeScript or JavaScript program |
Its return value and console output |
A catalog takes a typedrequest-shaped list as it is, without a tool() per method:
import type { IFlexToolCatalog } from '@modelprofile.com/flexharness';
// Build it once. FlexHarness prepares a catalog once per object, so return the same object every
// run and never change it afterwards; hand in a new object to change the catalog.
const catalog: IFlexToolCatalog<IProjectScope> = {
types: [
{ name: 'IMoney', schema: moneyJsonSchema },
{ name: 'IMovement', schema: movementJsonSchema },
],
methods: [
{
method: 'banking.listMovements',
summary: 'List the bank movements of an account.', // one line, at most 120 characters
description: 'Newest first. Paged with a cursor.',
inputSchema: {
type: 'object',
properties: { accountId: { type: 'string', description: 'The bank account.' } },
required: ['accountId'],
additionalProperties: false,
},
outputSchema: { type: 'array', items: { $ref: '#/$defs/IMovement' } },
examples: ['banking.listMovements({ accountId: "acc_1" })'],
async run(params, context) {
await context.provider.requestPermission({
kind: 'banking.read',
description: 'Read bank movements',
toolCallId: context.toolCallId,
});
return bankingService.listMovements(context.provider.scope, params);
},
},
],
};
const harness = new FlexHarness<IProjectScope>({
// ...
toolProvider: {
provideTools: () => ({ tools: eagerTools, catalog }),
},
});
run(params, context) receives the validated parameters and a context with the dotted method, the execute call's toolCallId, the run's abortSignal, and provider, the run's tool-provider context with its scope, session, run and requestPermission. It may return a value, a promise or an async iterable, exactly like a tool's execute. toolCatalog: { declarations: 'json-schema' } makes describe_tools return each method's description and self-contained JSON Schema instead of TypeScript.
Paths and search
A method path is up to 8 dot-separated JavaScript identifiers of letters, digits and _, at most 160 characters; JavaScript reserved words such as delete or import are refused, because a declaration cannot name them. A path cannot also be a namespace of another path. A host without namespaces has a flat catalog. search_tools ranks methods with BM25 over the last segment, its aliases, the namespace, the summary, tags and the description, after splitting camel case and stemming plurals; ties sort by path. An alias is another name of the method, such as one in its users' language (aliases: ['Beleg abrufen']), and weighs like the name. The catalog's vocabulary lists the words users say for the catalog's own words, { voucher: ['Beleg', 'Belege'], gross: ['Bruttobetrag'] }: a query that says every word of an alias searches for the catalog's words as well, so a German question finds the methods of an English catalog. Without query it lists a namespace: its child namespaces with their method counts, then its methods. Its description names the top-level namespaces with their sizes, so a first listing is often unnecessary. Results page with limit (default 10, at most 50) and offset. When one match clearly leads, the only match or one whose name's or alias's every word the query names, vocabulary words included, and that ranks at least 1.25 times higher than the next, the first page is followed by its declaration, as describe_tools would answer, so the model can call it without another step, as long as the whole answer fits the bytes describe_tools may answer; toolCatalog.describeTopMatch: false turns that off. On the lab's six tasks it saves six of 30 to 33 model steps and 15 to 16% of the prompt tokens, on 20 to 2,000 methods and a 330-method surface alike.
Declarations
describe_tools answers with TypeScript a model reads like an SDK:
// Declares types: IMoney, IMovement
/** An amount of money. */
interface IMoney {
/** Minor units. (integer) */
amount: number;
currency: "EUR" | "USD";
}
interface IMovement { id: string; amount: IMoney; bookedOn: string }
declare namespace banking {
/**
* List the bank movements of an account.
*
* Newest first. Paged with a cursor.
* @param params.accountId The bank account.
* @example banking.listMovements({ accountId: "acc_1" })
*/
function listMovements(params: { accountId: string }): Promise<IMovement[]>;
}
The summary and description form the JSDoc; every top-level parameter with a description or a constraint a type cannot state (integer, format, bounds, length, pattern, default) becomes an @param line, and examples become @example lines. Enums are literal unions, $refs to catalog types or to a schema's own $defs are named types, and outputSchema is the return type (Promise<unknown> without one). The same catalog and request always give byte-identical text: methods sort by path, types by name, and the layout depends only on the schemas.
Named types are declared once per session. A $ref to #/$defs/<Name> or #/definitions/<Name> that a schema does not define itself names the catalog type Name; a schema's own identifier-keyed definitions become named types as well, with a _2 suffix when two schemas define different shapes under one name. A type name that is no identifier, such as IEvent_order.created, is declared under one made of it (IEvent_order_created), and its declaration names the original. A declaration brings the named types its parameters use, transitively, minus those the model already has. Its result only names its types: toolCatalog.outputTypeDepth (default 0) sets how many levels of them it declares, and a // Not declared yet: line lists the rest, which describe_tools({ types }) declares when the model needs to read a result's fields. A deep result graph would otherwise outweigh the method many times: a voucher read whose result names 37 types costs 238 tokens instead of 8,154. FlexHarness reads that from the model's own context, the // Declares types: lines of earlier describe_tools and search_tools results the model can still see, so a compaction or an undo that removes a result makes its types appear again, and a restart changes nothing. describe_tools({ types: [...] }) declares types again on request.
A host whose source of truth is TypeScript hands in prebuilt text: declaration on a method (function listMovements(params: …): Promise<…>;, optionally after its JSDoc, named after the last path segment) or on a type (interface IMoney { … }). The model then reads that text; FlexHarness finds the catalog types it names by the identifiers of its code, its comments and string literals left out. A method declaration's parameter list brings its types in full and its result's types follow outputTypeDepth, as a generated declaration's do. FlexHarness still validates parameters against inputSchema, so a type a method's inputSchema refers to needs a schema as well. /typedrequest produces both for a typedrequest API (see Typedrequest APIs).
Calls, validation and transcript
execute checks params against the method's JSON Schema with @cfworker/json-schema (drafts 4, 7, 2019-09 and 2020-12 by $schema, 2020-12 without one), then with the schema's own validator when it has one, such as zod's, whose output reaches run. A failed call is a failed tool call whose error text is JSON the model can repair from:
{"error":"invalid_params","method":"banking.matchMovement","issues":[
{"path":"params.movementId","expected":"string","got":"5"},
{"path":"params.voucherIds","expected":"string[] (at least 1 item)","got":"[]"}],
"see":"describe_tools({\"methods\":[\"banking.matchMovement\"]})"}
An unknown path returns unknown_method with up to three didYouMean paths, or a hint to list it when it is a namespace. run executes inside the execute tool call, so the tool wrapper bounds its output with toolOutputLimits and projects a thrown error like any tool's. The run's tool part, its part.* events, the stored messages, FlexHarnessTranscript rows, and permission requests (through the part's toolCallId) carry the method path and its params, never execute. Only a call that names no method of the run, which fails with unknown_method, stays an execute part. A validator that throws instead of reporting a failure is a host failure: it is projected like an error run throws. The canonical Agent events keep the call as the model made it, execute with { method, params }, because that is what the model is shown again; listUncertainToolExecutions() reports those Agent records.
Session and agent allowedTools lists admit method paths like tool names; search_tools, describe_tools and execute follow automatically while at least one method is admitted, and the others do not exist for the run, nor do the named types only they use. While a run's tool provider returns a catalog, the three names are reserved and a method path may not equal a tool name. Resource tool providers cannot supply a catalog. FLEX_TOOL_CATALOG_LIMITS lists the bounds: 4,096 methods and types, 16 KiB of description text, 64 KiB of declaration text and per schema, 16 aliases of up to 64 characters per method or vocabulary word, 1,024 vocabulary words, 16 methods and 32 types per describe_tools call, and 64 KiB per answer (less when toolOutputLimits.maxBytes is lower); methods past the limit are named for another call.
Why one stable execute tool
Anthropic and OpenAI, including the ChatGPT Codex backend, place tool definitions before the messages in the cached prompt prefix. A tool set that changes mid-session rewrites that prefix at full price; a stable one stays cached. The catalog therefore sends the same three tools on every step and adds what the model learns to the conversation, where it is cached as it grows. Activating described methods as native tools for the following steps (activeTools) was left out: every activation would change the tool set and invalidate the whole cached prefix, and keeping activations across turns needs session state, so it is not cheap. Provider tool search (Anthropic's tool search with defer_loading, OpenAI's tool_search, both in the AI SDK 7 providers) is not used either: it is provider-specific, puts provider items into the canonical history that another model of the session cannot replay, ranks and describes tools without the catalog's summaries, namespaces and shared types, and is not established for the ChatGPT Codex backend.
FlexToolCatalog.create(catalog, options) prepares a catalog without a session, for checks and measurements: search(), describe(), prepareCall(), restrict(), and toolDefinitions(), the three tools exactly as the model receives them. Measured on the 300-method catalog of test/helpers.toolcatalog.ts, whose methods share six named types (o200k tokens of the request JSON; a provider's own rendering differs):
| What | Tokens |
|---|---|
| 300 eager tools, on every step | 104,051 |
| The three catalog tools, on every step | 412 |
| One eager input schema, median and maximum | 294 and 603 |
describe_tools for a median method, with its parameters' shared type |
245 |
The same method with its result's types declared too (outputTypeDepth: 32) |
404 |
| The same method once its types are declared | 163 |
| The same method as self-contained JSON Schema | 267 |
| Ten methods described one at a time: TypeScript with shared types, inlined, JSON Schema | 2,587, 5,259 and 3,703 |
A first call of a method costs two more model steps than an eager tool (search_tools, then describe_tools), one when the search's top match is clear; a known method costs none, and one describe_tools call covers up to 16 methods. test/test.tool-catalog.budget.node.ts keeps the three tools for 300 methods under 3 KiB of request JSON.
Proposals and denials
A method's onCall(params, context) decides each validated call before it runs: 'run' (the default without the hook), 'propose', or { deny: reason }. A denied call fails with {"error":"denied","method":…,"reason":…}. A proposed call does not run: it goes to the catalog's onProposals(proposals, context), whose return value is what the model sees, such as the ID of a confirmation card. execute hands over its one proposal; a run_code program's proposals are handed over as one batch once the program finished, and a program that fails proposes nothing. Each proposal carries its call's id, method and validated params.
const catalog: IFlexToolCatalog<IProjectScope> = {
onProposals: async (proposals, context) => cards.create(context.provider.scope, proposals),
methods: [{
method: 'vouchers.cancel',
summary: 'Cancel a voucher.',
inputSchema: cancelVoucherSchema,
onCall: () => 'propose',
run: async (params, context) => vouchers.cancel(context.provider.scope, params),
}],
};
Typedrequest APIs
/typedrequest makes the catalog of an API built on @api.global/typedrequest from its contracts. describeTypedRequests() reads them from their TypeScript sources with the optional peer @api.global/typedopenapi: every typed method becomes a method <area>.<method>, its area the file that declares it, with its declaration as the contract writes it and its JSDoc, and its request's and response's JSON Schema; the named types they reach bring their declarations and schemas. omit leaves request properties the host fills, such as the caller's identity, out of what the model sees and sends, and each method lists in omitted those its request has. The result is plain JSON, for a build step to write and a server to read.
import { catalogFromTypedRequests, describeTypedRequests } from '@modelprofile.com/flexharness/typedrequest';
const surface = await describeTypedRequests({
tsconfig: './tsconfig.json',
contracts: './ts_interfaces/requests/*.ts',
omit: ['identity'],
});
const catalog = catalogFromTypedRequests<IAppScope>(surface, {
router: server.typedrouter,
request: (method, context) => ({ identity: identityOf(context.provider.scope) }),
});
catalogFromTypedRequests() routes each call through the router's routeAndAddResponse() in process, its request the validated parameters with what request fills of the omitted properties the method's request has, and rejects a refusal with FlexTypedRequestError, the API's own error text and data. Deprecated methods stay out unless include lets them in, and the catalog's types are the ones the included methods reach, through their schemas' $refs and the names their declarations' code uses (FlexToolCatalog.identifiersIn(), which skips comments and string literals as the catalog does): a type only methods left out use is neither declared nor described. A kept type whose name is no identifier keeps the types that took the lower identifiers of its name (FlexToolCatalog.identifiersOf()), so it is declared as it is in the whole surface. aliases, vocabulary, onCall and onProposals work as for any catalog. readOnly marks the methods that only read, including their policy hook: onCall still decides every call, and allowed reads run concurrently (Tool Call Order). A policy that proposes changes must explicitly return run for the reads it allows. params makes the request of the model's parameters, such as IDs for the short names the model was shown, and response decides what the model sees of an answer, such as a response without a credential's secret. On a 330-method API, the catalog made from its sources takes the same model steps on the lab's six tasks as one whose declarations are generated from the JSON Schemas, and 1.7% (catalog mode) to 2.1% (code mode) more prompt tokens, the JSDoc its authors wrote. The subpath readme lists the details.
omitRequestPaths adds method-specific input omissions, for example { saveConnection: [['connection', 'password']] }; * traverses array/tuple elements. It requires @api.global/typedopenapi 0.10.0 or later when describing the surface. Schemas and declarations omit only the selected input leaves, retaining siblings, shared types and responses. request fills only those leaves without replacing public siblings or adding array entries, and omitted inputs are stripped before and after params. This policy is identical for execute and run_code; proposed changes invoke neither mapper, trusted fill nor router. Input omissions do not redact results: the host's response hook still owns secret removal from answers.
MCP tools
/mcp makes the catalog of MCP servers' tools with createMcpToolCatalog(). It takes the tools each server lists (createMcpTools()' servers are such a list) and a resolveClient that names the client a call runs on, so one catalog object serves every run while the host connects its clients per run:
import { createMcpToolCatalog, createMcpTools } from '@modelprofile.com/flexharness/mcp';
const connected = new Map<string, Awaited<ReturnType<typeof createMcpTools>>>();
// Once per configuration revision, from the tools the servers list: the `servers` of a
// createMcpTools({ servers: mcpServers }), each with its configuration's `trustReadOnlyHints`.
const { catalog } = createMcpToolCatalog<IAppScope>({
servers: listedServers,
prefixToolNames: true,
resolveClient: ({ serverName, call }) => connected.get(call.provider.runId)!.servers
.find((server) => server.name === serverName)!.client,
authorizeToolCall: (context) => context.call.provider.requestPermission({
kind: 'mcp.call',
description: `Call ${context.exposedToolName}`,
toolCallId: context.toolCallId,
}),
});
const toolProvider = {
async provideTools(context) {
const mcp = await createMcpTools({ servers: mcpServers, prefixToolNames: true });
connected.set(context.runId, mcp);
return {
tools: {},
catalog,
close: async () => {
connected.delete(context.runId);
await mcp.close();
},
};
},
};
A method's path is the name createMcpTools() gives the tool (agl__session_read), so allowlists, permissions and transcripts name it the same either way; a name that is no identifier, with a - or a reserved word, is declared as FlexToolCatalog.identifierOf() makes it, and toolNameMap maps each path to its server and tool. A prefixed method has the tool's own name as a search alias, so a query that names the tool without its server (session read, session_read) gets agl__session_read's declaration, as one naming the method does. Its summary is the first line of the tool's description, cut to 120 characters, and the description follows in full up to 16 KiB. authorizeToolCall decides execute and run_code calls as createMcpTools()' hook decides tool calls, before the client is resolved: it resolves to allow a call and throws to deny it, and call is the catalog's run context with the run's requestPermission. A method is read-only when its server's entry in servers sets trustReadOnlyHints: true and the server annotates the tool readOnlyHint: true, as createMcpTools()' isReadOnlyToolCall classifies it, or when isReadOnly declares it; no server's annotations are trusted by default. Its execute calls and the calls of a run_code program run concurrently with other reads (Tool Call Order), and each still passes authorizeToolCall: when that asks for permission, the requests of concurrent reads are pending together.
On AGL's 69 MCP tools (44 of its own, 25 of a Playwright server; o200k_base tokens, Anthropic request shape), the request of a one-line prompt takes 11,483 tokens with the tools as createMcpTools() offers them, and 470 with the catalog's fixed tools, or 647 with run_code, however many tools the servers list (85 take as many). A task that calls one tool reads 23,058 input tokens over its two steps as tools and 3,158 over four as a catalog: the model searches and describes a method before its first call, two steps that a clear top match, whose declaration the search attaches, shortens to one.
Code mode
toolCatalog: { codeMode: true } adds a fourth fixed tool, run_code, for work that takes several calls: the model writes a short TypeScript or JavaScript program, and only its return value and console output come back.
const ids = await api.vouchers.list({ month: '2026-09' });
const vouchers = await Promise.all(ids.map((voucherId: string) => api.vouchers.get({ voucherId })));
console.log('read', vouchers.length);
return vouchers.filter((voucher) => voucher.total.amount > 10000).map((voucher) => voucher.id);
Code mode needs the optional peers: pnpm add quickjs-emscripten-core @jitl/quickjs-ng-wasmfile-release-sync amaro. A run whose catalog has code mode fails at its start without them.
codeis the body of an async function.api.<dotted path>(params)calls a method; a refused call rejects with anApiErrorwhosedetailsare the JSONexecutewould report (invalid_params,unknown_method,denied, orfailedwith the projected message of an error the method threw).console.log,info,warn,erroranddebugare captured.- The answer is
{ result, calls, logs?, proposals? }: the returned JSON value, the number of calls, the console output, and whatonProposalsreturned. A program that does not finish is a failedrun_codecall whose error text is JSON:errorissyntax_error,program_error,limit_exceeded(withlimit) oraborted, withname,message, the program'slinewhere known,callsandlogs.nameandmessagekeep up to 2 KiB each, andtoolOutputLimitsbounds the whole report. - A program owns its calls and must await them. One that returns before each of its calls answered fails with
program_errorand proposes nothing. Its calls run while it computes, up tomaxConcurrentCallsat a time. When a program ends or stops, calls not started yet never start, and calls in flight see theirabortSignalabort and get up to 5 seconds to settle beforerun_codeanswers; a permission wait of such a call ends with it. - Every call a program makes is dispatched like an
execute: validation,onCall, the method'srunwith its permission requests,toolOutputLimits(an answer over the budget arrives as a result handle, see Large results), error projection. Read-only calls run concurrently and the others in the order the program made them (Tool Call Order). It is a tool part of its own, named by the method path, with the call ID<run_code call ID>.<n>andparentToolCallIdset to therun_codecall; its events, the stored messages andFlexHarnessTranscriptcarry it, and a permission request names it by that ID. A program's calls are recorded in at most half the roomcallbackLimitsleaves once the program's answer is set aside, so the run keeps the rest for its next steps: an output past that room is recorded cut, astoolOutputLimitscuts it, and an error message is shortened, while the program gets either whole, and a call that does not fit at all stops the program withlimit_exceededand limitcallbacks, as a failedrun_codecall of a run that goes on.FlexHarnessTranscriptrows stay flat for now: showing the calls inside theirrun_coderow needs a change in@design.estate/dees-catalog's harness components. The model's history holds only the program and its answer.
Limit (codeMode.limits) |
Default | Maximum | When it is exceeded |
|---|---|---|---|
maxSourceBytes |
32 KiB | 256 KiB | the program does not start |
cpuMs |
5 s | 60 s | the program stops; work between two of QuickJS's time checks, such as one long native call, runs at most 250 ms past it |
wallMs |
60 s | 30 min | the program stops; waiting for a permission decision does not count |
memoryBytes |
64 MiB | 512 MiB | the allocation fails with an InternalError the program can catch; uncaught, it stops the program |
stackBytes |
64 KiB | 256 KiB | deep recursion stops the program |
maxCalls |
200 | 10,000 | the program stops at the call past it |
maxConcurrentCalls |
8 | 64 | further calls wait for a slot |
maxLogBytes |
16 KiB | 256 KiB | further console output is cut off |
The sandbox is QuickJS-ng compiled to WebAssembly, through quickjs-emscripten-core. Programs run in worker threads, so the host goes on while one computes. The harness owns its workers: a run whose catalog has code mode starts one for its scope as it begins, so QuickJS loads while the model writes the program, and harness.dispose() stops the programs running and ends every worker. A worker runs one program at a time, each in a QuickJS runtime and context of its own that it frees once the program ended, so nothing of a program reaches the next: no globals, changed prototypes, pending jobs or late call answers. A worker serves the programs of one scope only, so programs of different scopes never share a WebAssembly instance. A program stopped at a limit, by an abort or by a failure of the host ends with its worker, which frees all its memory, and so does a worker that fails. A worker is replaced after 100 programs, or once its WebAssembly memory, which never shrinks, has grown past 64 MiB. Up to 4 workers wait for programs, each for at most 60 s, and a waiting worker does not keep the process alive. On a new worker a program takes about 65 ms, most of it starting the worker and loading QuickJS, and several hundred milliseconds where workers load a TypeScript loader such as tsx; on a waiting worker a trivial program takes about 1.5 ms and one with five calls about 2 ms. Each program sees only the language, api, ApiError and console: no require, import, process, network, files or timers. Values cross the boundary only as JSON. QuickJS runs on its worker's stack, from a shallow start, so the default stack limit trips QuickJS's own check before the worker's. Should the worker's stack still run out, the program stops at its stack limit; a failure of the worker itself fails the run_code call like an error a tool throws. QuickJS is not audited; FlexHarness does not rely on it alone: every call is validated, decided by onCall and permission-checked by the host as for execute. Types are stripped with amaro, the TypeScript stripper Node.js uses, so a program keeps its line numbers; type-only syntax that needs code (enum, namespace) is refused as a syntax error.
For 20 voucher lookups with a filter (test/test.code-mode.budget.node.ts), execute takes 3 model calls, reads 15,278 prompt bytes and grows the context by 14,428 bytes; run_code takes 2 model calls, reads 2,042 bytes and grows the context by 1,846 bytes.
Large results
In a run with a catalog, an output larger than toolOutputLimits.maxBytes is not truncated: FlexHarness stores it whole (up to resultLimits.maxStoredBytes, 8 MiB by default) in the session, and the model gets a handle with a summary of its shape and start:
{"$flexType":"result","handle":"res_5f0c9a1d2e3b4c67","bytes":61042,"pages":1,
"summary":{"type":"object","size":2,"keys":{"total":"number","books":"array(300)"}},
"read":"Too large to return whole: read it page by page with execute({ method: \"results.get\", params: { handle: \"res_5f0c9a1d2e3b4c67\", page: 1 } }), or in run_code with await api.results.get(\"res_5f0c9a1d2e3b4c67\", { page: 1 }); a dotted `path` reads a part of it."}
results.get({ handle, page?, path? }) is a built-in, read-only catalog method; search and namespace listings leave it out, since the handle names it. It answers { handle, path, page, pages, items | entries | chars, value }: one page of the value at path (a dotted path such as books.299.title), each page within the output budget, holding whole array items, whole object entries or a slice of a string, with the range it holds. A part too large for a page stands as {"$flexType":"oversize","path":…}, to read by its path. In code mode a program reads pages with api.results.get(handle, { page, path }); a method's answer over the budget reaches the program as a handle too. Unknown handles, paths and pages fail with unknown_handle, unknown_path and unknown_page. A handle reads only the session that stored it. A session keeps its latest results, 32 by default (resultLimits.maxResultsPerSession): storing one more deletes the oldest, and the rest are deleted with the session.
This applies to every tool of a run with a catalog, execute, run_code and the calls inside it included, but not to outputs that stream. Results need the stores' results domain, which the built-in stores provide; without it, or without a catalog, an output over the budget is truncated as before. A budget too small to hold even the handle keeps the truncated output.
The host option resultLimits bounds what a session stores, at most maxResultsPerSession results of at most maxStoredBytes JSON bytes each:
| Field | Default | Rule |
|---|---|---|
maxStoredBytes |
8 MiB (FLEX_RESULT_LIMITS), or toolOutputLimits.maxBytes if larger |
at least toolOutputLimits.maxBytes; a larger output is stored cut to it, and the cut parts stand as {"$flexType":"truncated",…} |
maxResultsPerSession |
32 | a positive integer; storing one more deletes the oldest |
By default a session stores up to 256 MiB. A host whose result store holds small documents sets both lower: resultLimits: { maxStoredBytes: 1024 * 1024, maxResultsPerSession: 16 } stores at most 16 MiB a session, 1 MiB a result.
Project Management Tools
Harness-owned project tools are opt-in and session-local:
builtInTools: {
renameSession: true,
projectManagement: {
task: true,
goal: true,
scratchpad: true,
},
},
renameSession enables rename_session. The projectManagement block enables the public project-management APIs and contains the model-tool flags; task, goal, and scratchpad each default to enabled unless explicitly set to false. Without that block, the public project-management APIs reject with FlexHarnessValidationError, while the required stores.projectManagement domain still participates in session cleanup. With no builtInTools configuration, none of these four tools is present. Constructor options are copied and frozen.
Project-management records use FLEX_PROJECT_MANAGEMENT_SCHEMA_VERSION, currently 1, and form a strict live-or-tombstone union:
interface IFlexProjectManagementSnapshot {
schemaVersion: 1;
revision: number;
sessionGenerationId: string;
sessionGenerationSequence: number;
goal?: string;
scratchpad: string;
tasks: Array<{
id: string;
content: string;
status: 'pending' | 'in_progress' | 'completed' | 'cancelled';
priority: 'high' | 'medium' | 'low';
createdAt: string;
updatedAt: string;
}>;
}
interface IFlexProjectManagementTombstone {
schemaVersion: 1;
revision: number;
sessionGenerationId: string;
sessionGenerationSequence: number;
deletedAt: string;
}
type TFlexProjectManagementRecord =
| IFlexProjectManagementSnapshot
| IFlexProjectManagementTombstone;
The tools use strict action-discriminated inputs:
task:list,create,update,delete, orclear. Create defaults topendingandmedium.goal:get,set, orclear.scratchpad:get,set,append, orclear. Append concatenates the supplied content exactly.rename_session: sets the active session title and returns the authoritative session.keep_cache_warm(cacheWarm: true, root sessions only): keeps the session's prompt cache warm; see Keeping the Prompt Cache Warm.
Every project action returns the authoritative revision and state; task mutations also return the affected task, and clear returns the removed tasks. Reads never save. A set, clear, append, update, idempotent create, or empty task clear that makes no state change returns the current revision without writing. Mutations load once, apply once, validate the complete next snapshot, and issue one compare-and-swap save at revision + 1. FlexHarness never retries or merges an external conflict.
Tool task creation accepts an optional id. When omitted, FlexHarness requires the stable AgentSession toolCallId and derives task_ plus the SHA-256 of JSON.stringify(['flexharness-project-task-v1', storageKey, sessionId, runId, toolCallId]). Repeating an explicit or deterministic ID with identical content, status, and priority is idempotent; different creation data conflicts. Application callers must supply an explicit id to createProjectTask() because no tool-call identity exists at that boundary.
The same engine is available to applications:
await harness.getProjectState(scopeId, sessionId);
await harness.listProjectTasks(scopeId, sessionId);
await harness.createProjectTask(scopeId, sessionId, { id, content, status, priority });
await harness.updateProjectTask(scopeId, sessionId, { id, content, status, priority });
await harness.deleteProjectTask(scopeId, sessionId, id);
await harness.clearProjectTasks(scopeId, sessionId);
await harness.getProjectGoal(scopeId, sessionId);
await harness.setProjectGoal(scopeId, sessionId, goal);
await harness.clearProjectGoal(scopeId, sessionId);
await harness.getProjectScratchpad(scopeId, sessionId);
await harness.setProjectScratchpad(scopeId, sessionId, content);
await harness.appendProjectScratchpad(scopeId, sessionId, content);
await harness.clearProjectScratchpad(scopeId, sessionId);
Public writes use { actor: 'application' }. Tool writes use { actor: 'agent', runId, toolCallId, agent? }, allowing custom stores to preserve attribution. Project side effects commit independently of the later model outcome and are intentionally outside transcript undo/redo.
FLEX_PROJECT_MANAGEMENT_LIMITS exports the hard UTF-8 and aggregate limits: goal 8 KiB, scratchpad 128 KiB, task content 8 KiB, task ID 512 bytes, title 2048 bytes, 512 tasks, and a 1 MiB serialized snapshot. The aggregate bound leaves room for worst-case JSON escaping of a controller-valid scratchpad. Loaded snapshots reject extra fields, duplicate IDs, invalid status/priority/timestamps, non-JSON data, wrong schema/revision, and every exceeded bound before use.
IFlexProjectManagementStore is exact per (storageKey, sessionId): load, CAS save, CAS tombstoneSession, and purgeNamespace must not collapse multiple sessions or storage namespaces. load() returns TFlexProjectManagementRecord | undefined and receives an optional IFlexProjectManagementSessionContext as its third argument; tombstoneSession() receives the same optional context as its fifth argument. FlexHarness always supplies both, while existing two-argument loads, four-argument tombstones, and shorter store implementations remain compatible.
IFlexProjectManagementSessionContext contains sessionGenerationId, sessionGenerationSequence, and optional subagent. The atomic IFlexSubagentProvenance block contains parentSessionId, parentSessionGenerationId, parentSessionGenerationSequence, originParentRunId, originParentToolCallId, agent, and actual session depth. Both the context and its separately cloned nested block are frozen.
Same-generation live saves use normal revision CAS, and a same-generation tombstone permanently rejects later live saves. A higher sessionGenerationSequence with a different sessionGenerationId may replace only an older tombstone using expected revision 0; it cannot replace a live record. This resets the PM revision for a recreated core session while stale saves and tombstones from older generations remain fenced. Deleting a recreated session that made no PM writes still replaces the prior-generation tombstone with a revision-1 tombstone for the new generation.
Every newly created core session exposes and persists a sessionGenerationId plus its monotonic sessionGenerationSequence. FlexHarness generates a strong random ID when sessionGenerationId is omitted. Applications may supply the ID to createSession() when they need to persist creation authority before dispatch; a supplied ID must be nonblank, contain no control characters, and fit within FLEX_SESSION_GENERATION_ID_MAX_BYTES (128 UTF-8 bytes). FlexHarness still assigns the sequence atomically. A legacy scope session without those fields is assigned a deterministic bounded ID derived from its immutable storageKey, sessionId, and createdAt; FlexHarness persists the repaired scope snapshot before accepting work. Grouped core deletion tombstones retain both fields after live metadata is removed.
Normal Flex session cleanup always waits in-flight local project operations, then loads and CAS-tombstones stores.projectManagement, regardless of whether PM tools are enabled in that harness. Child cleanup persists one complete IFlexSubagentProvenance block on its core tombstone before live metadata is removed. One cleanup invocation reuses the identical doubly frozen context object across bounded CAS retries; restart or a later cleanup invocation reconstructs a new frozen context from the persisted provenance. If a concurrent same-generation save wins first, cleanup reloads and retries; unresolved conflict or store failure retains the core Flex session cleanup tombstone for a later retry. The durable PM tombstone is not physically removed during normal session cleanup.
purgeNamespace(storageKey) is the explicit destructive reclamation operation and physically removes every live record and tombstone in that exact PM namespace. Applications may call it only after serializing every scope alias, preventing new admission, awaiting retireScope() on every harness owner, and deleting or purging the application-owned core scope namespace. retireScope() itself remains non-destructive and never calls purgeNamespace(). Purging PM first, purging only one alias, or racing a stale harness can remove the fence that makes session-generation reuse safe. NoSqlFlexHarnessStores.purgeNamespace(storageKey, { fence: true }) closes that race for the database bundle: it purges every domain in one transaction and refuses every later record creation in the namespace, so the storage key is never reused.
InMemoryFlexProjectManagementStore is the standalone in-memory implementation. InMemoryFlexHarnessStores and JsonFileFlexHarnessStores include projectManagement as a required bundle member. assertFlexProjectManagementSnapshot() validates live records, assertFlexProjectManagementTombstone() validates tombstones, and assertFlexProjectManagementRecord() validates the union. createEmptyFlexProjectManagementSnapshot(sessionGenerationId, sessionGenerationSequence) returns a revision-0 live state for the supplied current generation.
Subagents
subagents enables the harness-owned delegate, delegate_result, and delegate_send tools when at least one definition exists and the session depth is below maxSubagentDepth. A tool provider cannot replace any enabled built-in; global and agent tool allowlists still intersect to control model-facing exposure. Delegation is asynchronous by default: delegate returns { taskId, status: 'running', runId } after durable child admission, so the parent can continue independent work. delegate_result({ taskId, waitMs? }) reads live status and, once settled, bounded final text and model identity; its optional wait is limited to 240 seconds and observes parent cancellation. Set foreground: true to await the final answer in delegate itself. Children remain owned by the parent turn: Stop cancels its exact child runs, and the parent joins them before releasing authority and finalizing. A child result enters the parent's next model context as untrusted task data. When the parent's model ends its turn while background children still run, the turn waits for the next child to settle and takes its result in with another step (while steps remain), so a delegated result is never lost because the parent finished first.
The model calls it with this exact input shape:
interface IDelegateInput {
description: string;
prompt: string;
subagentType: string;
taskId?: string;
foreground?: boolean;
}
delegate_send({ taskId, prompt }) sends follow-up instructions to a running direct child owned by this exact parent turn. Both fields are required and extra fields are rejected; taskId is non-empty and at most 512 UTF-8 bytes, and prompt is non-empty and at most 64 KiB. It returns { steerId, queueId, runId } on admission, with a stable steer identity derived from the parent run and tool call. The child takes the input at its next inference boundary, without interrupting a running tool or permission wait. A child waiting to finish while its own descendants run can take the follow-up and continue inference; those descendants keep running. Use delegate_result to inspect progress and collect the final answer.
Controllers can use the same admission path explicitly:
await harness.sendSubagentInput(scopeId, parentSessionId, parentRunId, taskId, {
steerId: 'follow-up-1',
prompt: 'Prefer primary sources and include the publication date.',
});
Sending requires the exact active parent run, its currently delegated direct child, matching captured configuration and parent/child generations, and live writable ancestors. It does not authorize a root to skip levels, or infer ownership from immutable creation-origin metadata: a resumed child belongs to the current delegate invocation. Foreign, unowned, stale, cancelled, settled, or ending runs reject; sending never starts an idle generation, restarts a child, or resumes one. General public steerPrompt() still forbids child sessions. Child sends share root steering's duplicate rejection, pending count/byte bounds, accepted/applied events, durable user-message application, and terminal unapplied-input accounting.
Before creating or resuming a child, FlexHarness requests permission on the parent run with kind: 'subagent.start', the parent toolCallId, and separately bounded harness-owned agent/task metadata. Its metadata includes childSessionId, the exact ID reserved for this call: FlexHarness derives it deterministically when taskId is omitted and copies the supplied candidate when taskId is present. A resume also retains that candidate as taskId. This block is not truncated by toolOutputLimits. The reserved ID binds permission handling before child creation but does not prove that the child exists or is owned: applications must treat the later delegated admission context as authoritative. The controller answers through the normal permission APIs. This request has no rememberKey, so always is invalid; controllers use once or reject.
Each new invocation creates a durable child IFlexSession with immutable parentSessionId, origin parentRunId, origin parentToolCallId, agent, and depth. New public roots persist depth: 0; legacy schema-1 roots may omit it. These fields are harness-owned; public createSession() accepts sessionId, sessionGenerationId, configurationRef, and title. A child inherits its parent's configurationRef and cannot resume under a different reference. Child sessions reject direct prompt(), startPrompt(), enqueuePrompt(), and schedulePrompt() calls and run only through the delegate tool. The model and tool resolver contexts receive optional immutable parentSessionId and agent values so integrations can apply agent-specific model and tool policy. Child prompts use the definition's modelHint, system, and maxSteps.
When upgrading from 9.x, callers that require the previous synchronous delegate result must pass foreground: true. Tool providers must also leave the reserved delegate_result and delegate_send names available when subagents are enabled.
The parent tool part receives childSessionId as soon as the child is acquired, so a host can offer an authorized live transcript. A foreground call also adds the resolved child model to its part and returns bounded JSON:
{
taskId: 'subagent_...',
status: 'completed',
text: 'The child final answer, limited to 64 KiB.',
model: { provider: '...', model: '...', displayName: '...', variant: '...' },
}
Omitting taskId creates a deterministic child for the parent session, run, and tool call. The model-visible tool description and taskId schema state this creation rule directly. Repeating that same invocation does not create another child. If the deterministic child already has messages, FlexHarness reports an uncertain prior execution and never silently reruns it. This preserves AgentSession's durable parent tool intent as crash authority; controllers use listUncertainToolExecutions() and reconcileToolExecution() for uncertain parent calls.
Supplying taskId deliberately resumes an idle, live child from a later run of the same immutable parent session and the same configured agent. The caller must use the exact ID returned by an earlier completed delegate call; an unknown ID fails with safe corrective guidance and never creates a child under the supplied label. Resume starts a new child prompt while retaining the child's original parent run and tool-call origin. A child owned by another parent or agent, a deleted child, an active child, a same-run resume, or a second acquisition of the same child within one later parent run is rejected. Parent cancellation propagates only to the exact child run started by that delegate call.
delegatedRunAdmissionProvider optionally adds an application-owned admission lease around each internally delegated child run. The public contracts are IFlexDelegatedRunAdmissionProvider<TScope>, IFlexDelegatedRunAdmissionContext<TScope>, and IFlexDelegatedRunAdmissionLease. The provider is never called for root prompts or direct public prompt APIs. It receives a frozen context containing the resolved scopeId, exact captured scope, and storageKey; the exact child sessionId, session generation, queue, and run; the exact current parent session generation, queue, run, and delegate toolCallId; immutable originParentRunId and originParentToolCallId; the child agent and depth; and the child run AbortSignal. For taskId resume, the current parent queue/run/tool-call fields identify this delegate invocation, while the origin fields remain fixed to the invocation that created the durable child.
delegatedRunAdmissionProvider: {
async acquireDelegatedRunAdmission(context) {
const admission = await controller.acquireDelegatedRun({
child: {
sessionId: context.sessionId,
sessionGenerationId: context.sessionGenerationId,
sessionGenerationSequence: context.sessionGenerationSequence,
queueId: context.queueId,
runId: context.runId,
},
parent: {
sessionId: context.parentSessionId,
sessionGenerationId: context.parentSessionGenerationId,
sessionGenerationSequence: context.parentSessionGenerationSequence,
queueId: context.parentQueueId,
runId: context.parentRunId,
toolCallId: context.parentToolCallId,
originRunId: context.originParentRunId,
originToolCallId: context.originParentToolCallId,
},
signal: context.signal,
});
return {
close: () => admission.close(),
};
},
},
Acquisition completes before generation-side branch reversion, context compaction, model resolution, application or resource tool-provider callbacks, and model execution. Child session store and runtime initialization may already have occurred before acquisition. Providers must honor the supplied AbortSignal; an abort can settle the active child and parent without waiting for an acquisition that ignores cancellation, while FlexHarness retains ownership and closes any lease returned later.
For a normally acquired lease, FlexHarness gives close() an awaited attempt after AgentSession generation and before canonical accepted, rejected, or interrupted finalization and the terminal prompt.finished event. If an abort detaches an acquisition that ignores its signal, the run may settle before acquisition returns; FlexHarness retains that owner and closes any late lease. close() must be idempotent and safe to retry after rejection. A close failure prevents successful child acceptance and remains owned by the exact child session generation for retry by later exact-session deletion, scope retirement, or disposal. Retirement and disposal truthfully wait for late acquisition and lease cleanup.
Limits are validated and frozen at construction: at most 32 unique definitions; names are non-empty and at most 128 UTF-8 bytes; descriptions 2048 bytes; optional model hints 512 bytes; optional system prompts 64 KiB; and optional maxSteps a positive safe integer. maxSubagentDepth defaults to 6 and must be a positive safe integer at most 8. The root is depth 0, so the default permits six child levels (depths 1 through 6); delegation tools are absent at the configured maximum. maxSubagentCallsPerRun defaults to 32 and must be a positive safe integer at most 128. A call slot is consumed synchronously at the start of every schema-valid delegate execution, before semantic bounds, subagent type/depth validation, permission, or child work. Inputs rejected by the tool schema never start delegate execution and do not consume a slot. After successful semantic validation, the child ID candidate is reserved for the rest of the parent run, including after permission rejection or later failure. Permission rejection creates no child session. Omitting taskId reserves a deterministic new child ID; supplying taskId reserves that unverified candidate and attempts resume after permission only if it identifies a resumable child. Delegate descriptions are non-empty and at most 256 UTF-8 bytes, prompts non-empty and at most 64 KiB, subagent types at most 128 bytes, and task IDs at most 512 bytes.
Scope snapshots remain schema 1 and legacy child tombstones without subagent provenance continue to load. New child tombstones write the atomic provenance block and reject partial, parent-generation-mismatched, depth-mismatched, or duplicate-origin records. FlexHarness versions before this provenance addition reject that new optional key under their strict reader, so downgrading or mixing old readers with newly written scope snapshots is unsupported.
Session-Bound Configuration
An application can bind a root session to one immutable configuration revision at creation:
const harness = new FlexHarness({
scopeResolver,
modelResolver,
sessionConfigurationResolver: {
async resolveSessionConfiguration(context) {
const revision = await configurationAuthority.loadExact(context.configurationRef, {
storageKey: context.storageKey,
sessionId: context.sessionId,
sessionGenerationId: context.sessionGenerationId,
sessionGenerationSequence: context.sessionGenerationSequence,
signal: context.signal,
});
if (!revision) throw new Error('The pinned configuration is unavailable.');
return {
subagents: revision.subagents,
slashCommands: revision.slashCommands,
allowedTools: revision.allowedTools,
};
},
},
});
const session = await harness.createSession(scopeId, {
sessionId,
sessionGenerationId,
configurationRef: revisionId,
});
const commands = await harness.listSlashCommands(scopeId, session.sessionId);
const result = await harness.executeSessionSlashCommand(scopeId, session.sessionId, '/review');
configurationAuthority is application-owned and must retain historical revisions. configurationRef is an opaque, nonblank, well-formed string of at most FLEX_SESSION_CONFIGURATION_REF_MAX_BYTES (1024) UTF-8 bytes, without control characters. FlexHarness persists only the reference in the scope snapshot; the resolver's callbacks and registrations remain in memory. The reference cannot change during a session generation. A delegated child inherits it, and a stored child with a different reference from its parent is rejected during load. Reusing a session ID or caller-supplied generation ID does not reuse the old registry: the generation sequence is also part of resolution identity.
The frozen resolver context includes the resolved scopeId, captured scope, storageKey, exact session ID and generation ID and sequence, reference, agent and depth, and an AbortSignal. A child context also includes its parent's session ID and generation ID and sequence. FlexHarness validates, copies, and freezes the returned subagents and slashCommands using the same limits as constructor registrations. Concurrent users of one exact session generation share the resolved registry. Failed or cancelled resolutions can be retried. Deletion, retirement, and disposal abort pending resolution; a late provider result cannot replace the registry. A bound session fails through the sessionConfigurationResolver external-error source when its configuration cannot be resolved. Unbound sessions continue using constructor registrations.
The resolver may also return allowedTools, and each subagent definition may contain its own allowedTools. These are exact final names shown to the model: application tool names, Flex built-in names such as delegate and task, or resource names from createFlexResourceToolName(). Omission leaves that level unrestricted; an empty array exposes no tools. A bound root uses its resolved list. A delegated child uses the intersection of its resolved list and its named agent's list; a parent's list controls whether it can call delegate, while the child receives its own list. Unbound roots remain unrestricted, and unbound children may use constructor subagent lists. Each list is copied and frozen, with at most 128 unique, nonempty, well-formed names of at most 512 UTF-8 bytes each. Names are literal: globs and aliases are not interpreted, and a name absent from the current resource attachments grants nothing. FlexHarness applies the restriction after combining application, resource, and built-in tools; provider handles still close even if none of their tools are exposed. Tool allowlists limit availability and do not replace per-run permission decisions.
Use executeSessionSlashCommand() for bound sessions, including built-in commands. It resolves the session before custom command lookup. listSlashCommands() already selects the session registry. The existing executeSlashCommand() keeps its early parser and constructor-lookup behavior for unbound sessions; it rejects a known-command dispatch into a bound session. parseSlashCommand() remains synchronous and independent of configuration.
Instructions, application tools, and permission decisions remain per-run integrations through modelResolver, toolProvider, and the permission APIs. A host that supplies prompt or subagent system text must compose all intended instruction layers explicitly because that value overrides the model resolver's system. FlexHarness does not persist executable configuration or change those per-run APIs. Scope snapshots containing configurationRef cannot be opened by strict FlexHarness 9.6.1 readers; downgrade or mixed-version access to such snapshots is unsupported.
Sessions And Prompts
const session = await harness.createSession('project:billing', {
title: 'Invoice import',
});
const result = await harness.prompt(
'project:billing',
session.sessionId,
[
{ type: 'text', text: 'Extract the invoice totals.' },
{
type: 'file',
data: invoicePdfBase64,
mediaType: 'application/pdf',
name: 'invoice.pdf',
},
],
{ modelHint: 'document-model', maxSteps: 12 },
);
console.log(result.assistantMessage.parts);
console.log(result.usage);
TFlexPrompt is deliberately JSON-safe. It accepts a string or an ordered array of:
{ type: 'text', text }{ type: 'image', data, mediaType?, name?, sizeBytes? }{ type: 'file', data, mediaType, name?, sizeBytes? }
Attachment data is a string containing base64, a data URL, a remote URL, or a reference. Public input never requires Buffer or URL objects. Remote URL strings are converted only at the private AgentSession invocation boundary.
A reference is a URL of a scheme the harness's attachmentResolver reads, for files the host keeps itself, such as uploads too large to travel inline:
const harness = new FlexHarness({
// ...
attachmentResolver: {
schemes: ['exampleattachment'],
async resolve({ scopeId, scope, storageKey, sessionId, url, abortSignal }) {
const upload = await uploads.read(storageKey, sessionId, url.pathname, abortSignal);
return { data: upload.bytes, mediaType: upload.mediaType };
},
},
});
await harness.prompt(scopeId, sessionId, [
{ type: 'text', text: 'Book this receipt.' },
{ type: 'file', data: 'exampleattachment:3f2a', mediaType: 'application/pdf', name: 'receipt.pdf', sizeBytes: 48231 },
]);
The history keeps the reference, never the bytes. Each model call that sends it, the turns' and a compaction's summary call alike, asks the resolver for the bytes and sends them to the provider; the steps of one turn ask once per reference. The resolver receives the scope and session the reference is sent in, so it reads only that conversation's files, and a reference it cannot read fails the turn. Schemes are lowercase, without the colon, and cannot be data, file, http or https; the harness refuses a malformed resolver with FlexHarnessValidationError when it is made. sizeBytes, which a reference states because the harness cannot count its bytes, is accepted on references only; a URL of a scheme the resolver does not name is no reference and is sent as before.
Attachment payloads are never copied into public audit messages or events. Public attachment parts contain metadata only:
{
type: 'attachment',
partId: '...',
attachmentType: 'file',
source: 'inline-base64', // or data-url / remote-url / reference
sizeBytes: 48231, // omitted when it cannot be determined
mediaType: 'application/pdf',
name: 'invoice.pdf',
}
The original string remains only in canonical private Agent events, so a later model turn can receive the attachment again. A turn that succeeded stays in future context, and so does a turn stopped through abort() or cancelPrompt() once its model produced output: its prompt, the steers it took in and what the model produced before the stop, with every tool call that has no recorded result closed by an error result saying its effect is unknown. A turn stopped before any model output leaves no trace, so its prompt can be sent again without appearing twice; its run.finished event and finished queue entry carry keptInContext: false. Failed, resolver-failed, cleanup-failed, and persistence-failed turns, and turns a process restart interrupted, do not add anything to future context.
The main session methods are:
await harness.listSessions(scopeId);
await harness.createSession(scopeId, { sessionId, sessionGenerationId, configurationRef, title });
await harness.getSession(scopeId, sessionId);
await harness.getMessages(scopeId, sessionId);
await harness.listMessagePage(scopeId, sessionId, { limit: 50, before: cursor });
await harness.getMessage(scopeId, sessionId, messageId);
await harness.listSlashCommands(scopeId, sessionId);
const command = await harness.executeSlashCommand(scopeId, sessionId, '/init focus on tests');
const boundCommand = await harness.executeSessionSlashCommand(scopeId, boundSessionId, '/review');
if (command.type === 'prompt-admission') {
console.log(command.admission.queueId, command.admission.runId);
await command.admission.completion;
}
await harness.updateSession(scopeId, sessionId, { title: 'Renamed', archived: true });
await harness.updateSession(scopeId, sessionId, { title: null, archived: false });
await harness.getProjectState(scopeId, sessionId);
await harness.createProjectTask(scopeId, sessionId, { id: 'tests', content: 'Add tests' });
await harness.setProjectGoal(scopeId, sessionId, 'Ship the next release');
await harness.appendProjectScratchpad(scopeId, sessionId, 'One durable note.');
await harness.deleteSession(scopeId, sessionId);
await harness.deleteSessionGenerationCohort(scopeId, {
root: { sessionId, sessionGenerationId, sessionGenerationSequence },
authorizedCohort,
});
await harness.prompt(scopeId, sessionId, prompt, options);
const queued = await harness.enqueuePrompt(scopeId, sessionId, prompt, options);
console.log(queued.queueId);
await queued.completion;
const admission = await harness.startPrompt(scopeId, sessionId, prompt, options);
console.log(admission.queueId);
console.log(admission.runId);
await admission.completion;
const scheduled = await harness.schedulePrompt(
scopeId,
sessionId,
'refresh-index',
prompt,
{ debounceMs: 250 },
);
await harness.cancelScheduledPrompt(scopeId, sessionId, scheduled.scheduleKey);
await harness.getPromptQueueEntry(scopeId, sessionId, queued.queueId);
await harness.listPromptQueueEntries(scopeId, sessionId);
await harness.cancelPrompt(scopeId, sessionId, queued.queueId);
await harness.steerPrompt(scopeId, sessionId, admission.queueId, { steerId, prompt });
await harness.abort(scopeId, sessionId);
await harness.listPendingPermissions(scopeId, sessionId);
await harness.respondToPermission(scopeId, sessionId, permissionId, 'once');
await harness.pushRuntimeEvent(scopeId, sessionId, { type: 'workspace.changed', path: 'src/' });
await harness.listUncertainToolExecutions(scopeId, sessionId);
await harness.reconcileToolExecution(scopeId, sessionId, intentId, {
resolution: 'executed',
output: { committed: true },
});
const reversion = await harness.getSessionReversionInfo(scopeId, sessionId);
console.log(reversion.undoAvailable, reversion.redoAvailable, reversion.groups);
const undone = await harness.undoSession(scopeId, sessionId);
console.log(undone.revertedRunId);
const redone = await harness.redoSession(scopeId, sessionId);
console.log(redone.restoredRunId);
await harness.compactSession(scopeId, sessionId, { modelHint: 'summary-model' });
await harness.archiveSessionEvents(scopeId, sessionId, compactionEventId);
await harness.listBackgroundExecutions(scopeId, sessionId);
await harness.getBackgroundExecution(scopeId, sessionId, executionId);
await harness.abortBackgroundExecution(scopeId, sessionId, executionId);
await harness.retireScope(scopeId);
await harness.dispose();
Only one run may be active in a session. Additional prompts enter a bounded FIFO owned by FlexHarness, while different sessions can run concurrently. enqueuePrompt() resolves with { queueId, completion } after the immutable prompt and options have been accepted into that runtime queue and prompt.queued has been emitted. startPrompt() keeps its durable-admission behavior: it waits for its FIFO turn and resolves with { queueId, runId, completion } only after the canonical generation claim, run ID, and initial public audit messages have been durably reserved and the corresponding start events have been emitted. prompt() preserves the simpler behavior by awaiting completion internally.
Queue entries expose queued, starting, scheduled, running, completed, failed, and cancelled status through getPromptQueueEntry() and listPromptQueueEntries(). The list is ordered by process-local queueSequence. getPromptQueueEntry() throws FlexHarnessNotFoundError for an unknown or evicted ID. cancelPrompt() cancels one exact queue ID: a waiting entry leaves the FIFO and releases capacity immediately but remains queryable as cancelled until terminal retention evicts it; a promoted entry uses the canonical run cancellation path and is stopped with the same FlexHarnessAbortError ("The run was aborted.") as abort(); only a waiting entry is rejected with "The queued prompt was cancelled.". A cancelled run, whether stopped or found interrupted on reload, puts its error on the assistant message only; its user message is cancelled without an error, while a failed run marks both messages. Cancelling a terminal entry returns false, while an unknown ID throws. abort() remains scoped to the currently active run.
The displayed queue limits are the defaults. Outstanding count and byte limits apply per session and include every non-terminal queued or active prompt until it settles. Pending-admission limits apply to the complete harness while scope aliases are unresolved. Terminal retention applies per session. Exceeding an admission limit throws FlexHarnessQueueFullError.
Queue payloads, status records, and prompt.* queue events are process-local. The existing stores do not have a private generic queue domain: projections are deliberately redacted, Agent events are canonical conversation transactions, and jobs are AgentSession background executions. FlexHarness therefore never writes a never-started prompt into those unrelated domains. A process restart drops never-started entries; a prompt that reached durable run admission continues to use the existing canonical recovery policy and is repaired to a safe terminal state instead of being replayed.
schedulePrompt() waits for its FIFO turn, performs the same durable admission, exposes session status scheduled, and starts model preparation after its bounded debounceMs delay. Schedule keys remain unique across waiting and active prompts. cancelScheduledPrompt() returns true only while the matching schedule key can still be cancelled. Cancelling while it is still waiting rejects the schedulePrompt() call itself; cancelling after durable admission rejects the returned completion and marks its reserved audit messages cancelled.
The reservation save is the admission point. A save failure produces no start events or active audit. If disposal begins while that save is in flight and the save commits, admission still resolves and its completion settles as cancelled; disposal waits for terminal finalization.
listMessagePage() returns the newest contiguous page in chronological order. limit must be an integer from 1 through 50 and defaults to 50. nextCursor is opaque, limited to 4096 UTF-8 bytes, bound to the resolved storage namespace and session, and remains stable when newer messages are appended. Mismatched and stale cursors fail validation. Each page carries startIndex, the zero-based index of its first message in the session's visible history (for an empty page, the index the page ends at), so a consumer can order and page by absolute position. Appending messages leaves every index unchanged; a branch after an undo reuses the positions of the history it replaces, so a consumer re-reads its pages after session.history.changed. getMessage() performs an exact lookup. Transfer identifiers are limited to 512 bytes, text and reasoning parts to 96 KiB, complete messages to 480 KiB, and complete page envelopes to 512 KiB. A page may therefore contain fewer messages than requested. Oversized text is truncated and an otherwise oversized parts collection is replaced with an explicit elision marker; metadata that still cannot fit fails validation. Canonical private Agent events are unchanged.
updateSession() supports title replacement, explicit title clearing with null, and archive state through archived. Title-only updates remain available while prompts are queued or running, while permission is pending, and after archival. Requests containing archived are rejected while the session has any outstanding prompt or pending permission; a mixed title-and-archive request is rejected atomically without changing the title. Archived sessions expose archivedAt. Every session exposes preview once a run with prompt text was admitted: its first prompt as one line (whitespace runs collapsed, cut at a character boundary to maxTitleBytes), with no model call, so a host can name a session nobody titled (title || preview). It is set once, a session stored before previews existed takes it from its first prompt at its next run, and no caller sets or clears it. Deleting a session cascades through its complete descendant subtree. One durable root-keyed tombstone group hides every newly affected live session, and the delete also joins any already-separate descendant cleanup groups without rewriting their roots. FlexHarness then cancels queued and active subtree work, emits terminal queue events, waits for admitted initialization, and purges runtime queue status while cleaning runtime and persisted domains child-first. The requested root tombstone is removed last after every domain confirms cleanup; project-management cleanup confirmation is a retained durable project tombstone rather than physical removal. An imported session can hold no project-management state, so its deletion writes no project tombstone. A successful live deleteSession() call emits session.deleted for each session it newly tombstoned; retries of an existing tombstone and automatic load, retirement, or disposal cleanup emit no deletion events. Direct deletion of a descendant cascades only through that descendant's subtree. Cleanup authority follows the resolved storage namespace, so scope aliases share the same groups. A partial failure retains durable ownership for retry by a later deleteSession(), namespace load, retireScope(), or dispose() call.
deleteSessionGenerationCohort() is the generation-fenced destructive form for controllers. Its root and every authorizedCohort entry use the exact { sessionId, sessionGenerationId, sessionGenerationSequence } shape. The cohort contains unique entries in strictly ascending sessionId order and is limited by FLEX_SESSION_GENERATION_COHORT_MAX_ENTRIES (2048). FlexHarness atomically validates the matching root, its complete live subtree, and every retained descendant tombstone group before reserving deletion. A missing or different root generation returns { matched: false } without touching a newer generation; a missing or mismatched cascade entry throws before new tombstoning or cleanup. Extra valid entries do not expand the cascade. A matched partial cleanup remains retryable with the same authority and resolves to { matched: true } when cleanup completes. The existing deleteSession() method remains available for callers that intentionally own the complete reusable-ID namespace.
abort() returns true only while cancellation is still accepted. Terminal persistence is the run's commit point; once it starts, abort() returns false and the already-fixed terminal outcome completes while the session remains busy.
Steering a Running Turn
steerPrompt() hands a message to the running run of a prompt instead of queueing it behind that run:
const admission = await harness.startPrompt(scopeId, sessionId, 'Plan the release.');
await harness.steerPrompt(scopeId, sessionId, admission.queueId, {
steerId: 'steer-1',
prompt: 'Keep the database migration out of it.',
});
const result = await admission.completion;
console.log(result.unappliedSteerIds); // [] once the run took the steer in
The run takes a steer in at its next step boundary: after the tool results of its current model step are recorded and before its next model call. A running tool call or a pending permission is never interrupted; the steer waits for the boundary that follows it. A steer taken in is a user message of the run (steerId set) that the model sees in its next call and every later turn. If the run had already answered, the assistant message that answered is completed and the run answers on in a new assistant message after the steer; a steer taken in before the run answered anything is placed before its still empty assistant message. IFlexPromptResult.userMessage stays the prompt and assistantMessage is the run's last assistant message. Undo and redo treat the run's messages as one turn.
steerPrompt() resolves with { steerId, queueId, runId } once the run accepted the steer and prompt.steer.accepted was emitted; prompt.steer.applied carries the messageId of the user message the steer became. Every accepted steer ends in exactly one of two ways: it is in the model context of later turns, or it is listed in unappliedSteerIds, in arrival order, on IFlexPromptResult and run.finished, and on the finished queue entry when that list is not empty. A steer is listed when the run did not take it in before it ended, because its last model step was already running or because it was stopped or failed, and also when the run itself does not stay in context. keptInContext on run.finished and on the finished queue entry says whether it does: it is false for a failed run and for a run stopped before any model output, and such a run lists every steer it accepted, including the ones it took in and reported with prompt.steer.applied. Send its prompt and its unapplied steers again to have them answered. The run stops taking steers atomically when its generation ends or abort()/cancelPrompt() stops it. A steer for a prompt without a running run, or whose run is ending, is refused with FlexHarnessSteerRejectedError reason not-running; the caller sends it as the next turn instead. Reusing a steerId within a run is refused with reason duplicate, an unknown queueId throws FlexHarnessNotFoundError, and subagent sessions are not steerable. A run holds at most maxOutstandingPromptsPerSession pending steers, and their bytes count against the session's outstanding prompt byte limit until the run takes them in or ends, so startPrompt() and enqueuePrompt() are refused with FlexHarnessQueueFullError while pending steers fill it. They are process-local like queue entries: a process restart ends the run and drops its pending steers, while steers it took in are durable parts of the turn.
Keeping the Prompt Cache Warm
A provider's prompt cache expires when nothing reads it for its TTL, and the next turn then writes the whole prefix again. keepCacheWarm() keeps a root session's cache warm until a deadline:
await harness.keepCacheWarm(scopeId, sessionId, { durationMs: 60 * 60 * 1000 }); // warm until 1 h from now
await harness.keepCacheWarm(scopeId, sessionId, { durationMs: 0 }); // stop
const state = await harness.getCacheWarm(scopeId, sessionId); // { sessionId, warmUntil?, nextWarmAt?, idleMs?, basis? }
Each call replaces the deadline with durationMs from now, capped at cacheWarm.maxDurationMs (default 4 h, at most 24 h); 0 stops it. While the deadline has not passed and the session is idle, the harness sends a cache-warm prompt once the session has had no model call for the cache TTL less one minute: 4 minutes for a 5-minute cache, 29 for a 30-minute one, 59 for a 1-hour one. The TTL comes from the request the provider actually sent last, its basis in getCacheWarm():
basis |
TTL | Sent request |
|---|---|---|
anthropic-5m, anthropic-1h |
5 min, 1 h | Anthropic: the shortest cache_control lifetime of the request |
openai-30m |
30 min | OpenAI's API, GPT-5.6 and later: prompt_cache_options.ttl is 30 minutes after the last write or reuse |
openai-24h |
30 min | OpenAI's API, an earlier model with prompt_cache_retention: '24h' (cache: { retention: '24h' }): OpenAI keeps the prefix typically around 30 minutes and up to 24 hours, without promising more |
openai-in-memory |
5 min | OpenAI's API, an earlier model with in_memory or no retention; or the ChatGPT backend of ChatGPT accounts, which takes no retention and documents none. In-memory entries last 5 to 10 minutes of inactivity |
configured |
cache.warmDurationMs |
any |
With no basis (another provider or endpoint, or a custom transport that may change the request) the TTL is 5 minutes. No OpenAI retention promises a hit beyond about 30 minutes, so none makes warming unnecessary; extended retention only spaces warm turns further apart. Source: OpenAI's prompt caching guide, read 2026-10-11. A prompt, a steer or a running turn resets the wait, and no warm prompt is sent while a turn runs, a prompt is queued, an undo's redo history is open, a reversion is pending or a slash command runs. The warm prompt is the fixed text FLEX_CACHE_WARM_PROMPT, starting with [cache-warm], which asks for the single word ok. Its turn uses the model and system of the session's last turn and offers the same tools, so it reads the cached prefix, but it refuses every tool call, takes at most two steps, caps each model call at cacheWarm.maxOutputTokens (default 1024) and takes no steer: steerPrompt() refuses one with reason not-running. Keeping a session warm costs a cache read of its prefix and a few new tokens per warm turn; the warm exchanges stay in the model context.
Warm turns are marked so hosts can leave them out of what they show and of what they count as work: both messages, the session activity, the prompt queue entry, run.started, run.finished and model.usage carry cacheWarm: true. A warm turn changes neither the session's updatedAt nor its preview, and adds no undo step. Three events report the keeper: cache.warm.changed (source host or agent, warmUntil absent once stopped), cache.warm.completed with the warm turn's usage, whose cacheReadTokens and cacheWriteTokens are the provider's own counts and show whether it hit the cache (OpenAI's input_tokens_details.cached_tokens and cache_write_tokens, Anthropic's cache_read_input_tokens and cache_creation_input_tokens), its idleForMs since the session's previous model call and its basis, and cache.warm.stopped with reason deadline, cleared, failed (the warm turn failed; error set) or unavailable (the session or its scope went away).
builtInTools: { cacheWarm: true } gives root sessions the keep_cache_warm tool, so the agent can keep its own cache warm: { duration: '1h' } sets the deadline (amounts of s, m or h), 'off' or '0' stops it. The keeper is process-local: it ends with dispose() and does not survive a restart. It warms only the loaded scope it was set in: retiring the scope or deleting the session stops it with unavailable, and it never loads a scope again. Tests pass cacheWarm.clock (now, setTimeout, clearTimeout) to drive it with a fake clock.
Slash Commands
parseSlashCommand() is the public strict pure parser. It accepts at most 768 KiB, returns not-command for input not starting with /, malformed for invalid slash syntax, and parsed with the exact input, lowercase command name, separator-stripped raw argument text, and OpenCode-compatible quoted tokenization. Command names match [a-z][a-z0-9_-]{0,63}. The public isValidSlashCommandName(name) and isReservedSlashCommandName(name) helpers let applications validate registration names against the same rules. Single and double quotes group tokens and are stripped; escapes are not interpreted.
Applications register immutable custom commands at construction:
const harness = new FlexHarness({
// scopeResolver, modelResolver, and other options...
slashCommands: [
{
name: 'review-area',
description: 'Review one area of the workspace.',
template: 'Review $1 with these additional constraints: $ARGUMENTS',
},
{
name: 'refresh-index',
description: 'Refresh the application-owned workspace index.',
async handler({
scopeId,
scope,
storageKey,
sessionId,
sessionGenerationId,
sessionGenerationSequence,
rawArguments,
arguments,
signal,
}) {
return indexer.refresh({
scopeId,
scope,
storageKey,
sessionId,
sessionGenerationId,
sessionGenerationSequence,
rawArguments,
arguments,
signal,
});
},
},
],
});
compact, init, undo, and redo are reserved. listSlashCommands() verifies the scope and session and returns immutable data-only descriptors with kind, placeholder hints, current immediate availability, and workspaceReversion. compact, undo, redo, and custom handlers require an otherwise idle command session. Prompt templates and init use the normal bounded FIFO and may wait behind an active prompt. Their descriptors report availability from the same queue-admission conditions used by execution: lifecycle, slash ownership, pending reversion, root-session eligibility, outstanding count, and estimated prompt bytes. A dynamic capacity race may still produce FlexHarnessQueueFullError during admission. Only one slash-command execution may own a session at a time.
At most 128 custom commands may be registered. Every registration must be a plain object with exactly one of template or handler; names must match [a-z][a-z0-9_-]{0,63}, be unique, and not use a reserved name. Optional descriptions must be non-empty and at most 2048 UTF-8 bytes. Templates must be non-empty and at most 768 KiB, and the expanded prompt must also fit 768 KiB. Registrations are copied and frozen during construction.
/undo and /redo, plus undoSession() and redoSession(), move a durable history cursor. The direct methods return { revertedRunId } and { restoredRunId }; the slash forms return { type: 'operation', name: 'undo' | 'redo' }. Each committed cursor move emits one session.history.changed event with direction, runId, and the selected session identity. A committed branch emits the same event with direction: 'branch' and no runId, so controllers should refresh the complete selected session. Capture finalization, cleanup, and metadata-only changes do not emit this event.
A completed root-session turn defines an operation-group boundary. Failed and cancelled turns after it belong to that group; leading failed or cancelled turns belong to the first completed group. No completed boundary means there is nothing to undo. Undo applies selected segments in reverse order and redo applies them in forward order. Without turnReversionProvider, only transcript and future model context move. Hidden messages disappear from getMessages(), message pages, exact message lookup, and future model context. Starting a new prompt, template, handler, or compaction from an undone position commits a branch: hidden messages and segments are removed durably and cannot be redone. Successful event archival also commits hidden redo history; a missing compaction or failed archive leaves it intact.
Two horizons bound undo. Retention pruning removes the oldest complete visible units when reversionLimits is exceeded. Explicit event archival marks covered turns context-unavailable and prunes complete prefixes that can no longer be rebuilt; FlexHarness never crosses that archive horizon. Manual compaction without archival retains the original events and remains undoable. Schema-1 projection history and sessions migrated from 2.x have no reversion segments, so historical turns are not retroactively undoable; newly written turns are tracked normally.
Session metadata archival through updateSession(..., { archived: true }) only sets archivedAt. It does not archive Agent events, retire captures, or remove undo history.
executeSlashCommand() is the parser and constructor-command lookup boundary for unbound sessions; executeSessionSlashCommand() resolves bound sessions before lookup. Their result distinguishes not-command, malformed, unknown, completed operation, bounded handler-result, and prompt-admission. A prompt admission contains the normal { queueId, runId, completion }; await admission.completion for the model result. Unknown commands are never admitted as literal prompts. Known unavailable commands and invalid arguments throw typed FlexHarness errors. Options accept modelHint, system, maxSteps, and signal; commands do not accept attachments. /compact takes modelHint and signal only and refuses system and maxSteps with FlexHarnessValidationError. Aborting a template or init execution cancels its exact queued or started prompt without affecting another queue entry.
Templates replace every $ARGUMENTS with untouched raw argument text. $1 through the highest referenced positional placeholder use tokenized arguments, with the highest position receiving all remaining tokens joined by spaces. Missing positions become empty. A template with no placeholders appends non-empty raw arguments after a blank line. /init uses the OpenCode 1.18.15 AGENTS.md initialization prompt with provider-neutral active-workspace wording.
Handler context is frozen and contains only the resolved scope identity, session identity, required canonical sessionGenerationId and sessionGenerationSequence, raw and tokenized arguments, and an AbortSignal. Handler results are converted with the configured toolOutputLimits; void becomes JSON null. Handler failures use externalErrorProjector with source slashCommand. Same-session command overlap is rejected, including reentry from a handler. Scope retirement and disposal abort and await active handlers; prompt-admission commands transfer immediately to the normal prompt lifecycle.
Workspace Reversion Provider
reversionPolicy defaults to transcript-optional, preserving the V1 behavior described above. Set it to workspace-required when transcript and workspace traversal must move together. This policy requires an IFlexTurnReversionProviderV2 at construction.
Applications using the original protocol can continue to provide all six unchanged IFlexTurnReversionProvider operations:
import type { IFlexTurnReversionProvider } from '@modelprofile.com/flexharness';
const turnReversionProvider: IFlexTurnReversionProvider<IProjectScope> = {
prepare: (context) => workspaceSnapshots.prepare(context),
inspectCapture: (context) => workspaceSnapshots.inspectCapture(context),
finalize: (context) => workspaceSnapshots.finalize(context),
inspectApply: (context) => workspaceSnapshots.inspectApply(context),
apply: (context) => workspaceSnapshots.apply(context),
release: (context) => workspaceSnapshots.release(context),
};
Protocol 2 adds the protocolVersion discriminant and a tagged finalized outcome. The prepare, apply, apply-inspection, and release contexts remain the V1 shapes:
import type {
IFlexTurnReversionProviderV2,
} from '@modelprofile.com/flexharness';
const turnReversionProvider: IFlexTurnReversionProviderV2<IProjectScope> = {
protocolVersion: 2,
prepare: (context) => workspaceHistory.prepare(context),
inspectCapture: (context) => workspaceHistory.inspectCapture(context),
async finalize(context) {
const capture = await workspaceHistory.finalize(context);
if (capture.changedPaths.length === 0) {
return {
disposition: 'no-change',
reference: capture.cleanupReference,
};
}
if (!capture.revertible) {
return {
disposition: 'nonrevertible',
reference: capture.cleanupReference,
reasonCode: 'git.unmerged',
affectedWorkspaces: [{ id: capture.workspaceId, label: capture.workspaceLabel }],
};
}
return {
disposition: 'revertible',
reference: capture.reference,
affectedWorkspaces: [{ id: capture.workspaceId, label: capture.workspaceLabel }],
};
},
inspectApply: (context) => workspaceHistory.inspectApply(context),
apply: (context) => workspaceHistory.apply(context),
release: (context) => workspaceHistory.release(context),
};
const harness = new FlexHarness<IProjectScope>({
// scopeResolver, modelResolver, stores, and other options...
turnReversionProvider,
reversionPolicy: 'workspace-required',
});
V1 finalize() returns a JSON reference, and a finalized V1 inspectCapture() result is { status: 'finalized', reference }. V2 finalize() returns revertible, no-change, or nonrevertible, and a finalized V2 inspection is { status: 'finalized', outcome } with the same complete tagged outcome. Every V2 outcome carries a normalized cleanup reference. A revertible outcome also carries affectedWorkspaces. A no-change outcome omits that list or supplies an empty list. A nonrevertible outcome carries a stable reasonCode and may carry affected workspaces.
Affected workspace descriptors contain only stable, non-whitespace id and display label strings. A result accepts at most 64 unique descriptors; IDs are limited to 512 UTF-8 bytes, labels to 2048 bytes, and reason codes to 128 bytes matching [A-Za-z0-9][A-Za-z0-9._-]*. References use the configured JSON normalization depth and byte limit with an absolute 256 KiB cap.
Controllers query safe history metadata with one method:
const info = await harness.getSessionReversionInfo(scopeId, sessionId);
for (const group of info.groups) {
console.log(
group.runId,
group.kind,
group.visibility,
group.affectedWorkspaces,
group.affectedWorkspacesTruncated,
group.barrierCauses,
group.reasonCodes,
);
}
The immutable result exposes undoAvailable, redoAvailable, and groups classified as candidate, barrier, or no-change. Group metadata contains at most 64 unique affected workspaces; affectedWorkspacesTruncated is true when additional unique descriptors were omitted. barrierCauses states why a barrier blocks traversal, each cause once in this order: pending while a capture runs or its outcome is not final, nonrevertible for a workspace outcome that cannot be reverted, transcript-only for legacy transcript history without a workspace capture; it is empty for candidate and no-change groups (FLEX_SESSION_REVERSION_BARRIER_CAUSES lists the order). reasonCodes lists the distinct reasonCodes of the group's nonrevertible segments in segment order, so an application can say why a turn is not revertible even when its provider named no affected workspace. It never includes capture IDs or provider references.
workspaceSnapshots is application-owned. For the local filesystem tools, /tools/node provides InMemoryWorkspaceReversion, a protocol 2 provider that keeps each turn's file images in memory; its readme describes the observer it takes changes from. Every context contains scopeId, scope, storageKey, sessionId, the required canonical sessionGenerationId and sessionGenerationSequence, runId, deterministic captureId, and an AbortSignal. Apply contexts additionally contain the normalized reference, per-segment operationId, and direction; release contexts contain the reference. Deleted-session recovery retains the same generation from the durable tombstone, so providers can reject callbacks from an older same-ID generation.
The protocol is durable and inspectable:
- FlexHarness persists a
preparingcapture intent before callingprepare(), before model or tool execution. The provider must establish exclusive capture ownership for thatstorageKeyand retain it untilrelease()succeeds. inspectCapture()returnsmissing,prepared,finalized, orunknown. A finalized V1 result includesreference; a finalized V2 result includes the complete taggedoutcome.missingis safe only while the durable state is stillpreparing; a missing prepared or finalizing capture, anunknownresult, or an inspection failure fences the namespace.finalize()closes the capture and returns its JSON-safe V1 reference or complete V2 outcome. FlexHarness may call it during normal finalization or recovery afterinspectCapture()reportsprepared. Failed and cancelled root turns are captured too. A root capture spans its admitted subagent effects, including asynchronous children joined before finalization, although child transcript records remain separate. A failed parent cancels its outstanding children before joining them.- Before undo or redo, FlexHarness persists an apply write-ahead record.
inspectApply()must report the exact durable outcome for the suppliedoperationId:not-appliedmeans no effect occurred,appliedmeans the complete effect occurred, andunknownmeans the provider cannot prove either result. FlexHarness callsapply()only fornot-applied; after an apply error it inspects again, and anunknownresult or inspection failure fences the namespace. Progress is persisted after each segment, and the transcript cursor moves only after the complete unit succeeds. Known zero-progress failures leave the cursor unchanged and can be retried; partial progress retains the write-ahead record and resumes after restart. Caller cancellation is honored until the first workspace segment makes progress; recovery then continues with fresh bounded maintenance signals until the unit and cursor commit. EachoperationIdis unique to one apply of one segment: a later undo or redo gets a new ID even when it moves the same turn between the same cursors, such as undoing a turn again after a redo. Retries of that apply, in-process or after a restart, reuse its persisted ID, so providers must treatoperationIdidempotently and may record applied IDs to answerinspectApply(). release()relinquishes the capture after no-change or nonrevertible V2 finalization, branch commitment, retention or archive pruning, session deletion, or other durable removal. It must be idempotent: an unacknowledged release remains persisted and is retried before FlexHarness discards the reference.
Under workspace-required, a group is a barrier when any segment is pending, nonrevertible, or legacy transcript-only history. It is a candidate when at least one segment is revertible and none is a barrier; otherwise it is no-change. Only candidates can be traversed. No-change groups after a candidate travel with that candidate until the next candidate or barrier. Leading no-change groups remain visible. A retained barrier blocks older groups, while later candidates remain undoable. A mixed revertible/nonrevertible group is a barrier.
Inspection, recovery, finalization, and release use fresh maintenance signals bounded by the top-level reversionMaintenanceTimeoutMs option. It must be a positive safe integer no greater than 30 minutes. When omitted, FlexHarness uses agentSessionPolicy.generationLeaseCleanupTimeoutMs as a compatibility fallback, then defaults to 30 seconds when neither option is supplied. The agent-session setting continues to govern AgentSession generation-lease cleanup independently. Providers must observe every supplied signal and must serialize ownership for a storage namespace. A provider with the matching persisted protocolVersion must remain configured whenever a capture-backed session is reopened, retired, disposed, or deleted. Capture-backed recovery and deletion fail closed without it.
Workspace reversion is generic and application-defined. It does not reverse network, database, billing, or other side effects unless the provider deliberately captures them. References are normalized with toolOutputLimits and have an absolute 256 KiB encoded cap.
reversionLimits field |
Default | Hard maximum |
|---|---|---|
maxCompletedTurns |
100 | 1000 |
maxSegments |
300 | 3000 |
maxExcludedRunIds |
1000 | 10000 |
maxPendingReversionReleases |
1000 | 10000 |
Every configured value must be an integer from 1 through its hard maximum. Pruning removes complete prefixes rather than splitting an undo unit. Excluded run IDs prevent a committed branch from re-entering model context, while pending releases retain provider ownership until acknowledgement.
retireScope() stops runtime ownership for the complete resolved storage namespace without deleting its durable snapshot. It does not load a namespace that has no cached or in-flight state. For loaded state, it preserves and waits for persistence that has already started, while later queued reads, writes, and run admissions reject with FlexHarnessAbortError. It cancels queued prompts and cancellable runs, rejects pending permissions, emits queue terminal events, waits for committing runs, terminal persistence, queue drains, tool-handle closure, and detached tool-provider cleanup, then purges queue status and clears and evicts the cached state. Failed cleanup ownership remains cached so a later retireScope() or dispose() call can retry it. Calls through storage-key aliases share the same retirement drain. A later call can load the durable namespace again after successful retirement if the application still resolves it. A loaded namespace keeps every session's messages in memory, so retireScope() is also the supported way for a host to evict a scope it has stopped using, such as an opened archive of imported sessions. Retirement releases every staged session import and cancels queued prompts and runs, so a host retires only a scope with no open import and no work it still wants.
Normal retirement-induced cancellation does not make retireScope() reject. Unexpected failures observed through run finalization or scoped cleanup are surfaced without dropping the resources that still require cleanup. One such failure is thrown directly; multiple failures are reported through FlexHarnessRunError. Calling retirement or disposal again retries retained cleanup ownership.
Applications removing a scope must stop and serialize new admission across every alias before calling retireScope(), await retirement, and only then remove or purge application-owned durable records. FlexHarness cannot discover aliases before the application resolver returns. Integrations must not use retirement itself as durable deletion.
History And Audit Behavior
The model context is built by the session's canonical AgentSession event history as:
- Previous canonically accepted generations.
- The normalized current user message.
- AgentSession's result messages.
A failed or cancelled prompt remains visible through getMessages(), with failed or cancelled status, but is not included in future model context. Canonical Agent events remain private and are not exposed by the session or message APIs.
Resolved model identity contains provider and model IDs plus optional display name and effective variant. The identity, including its variant, is attached to a failed assistant message when resolution completed before a later failure, matching the provider/model behavior. Prompt results contain it only on success.
Public audit history is safe to send to controllers: attachment parts contain source and size metadata, never inline base64, data URLs, or remote URL payloads. Canonical private Agent events retain those values solely for subsequent model turns.
Sessions expose idle, scheduled, running, waiting_permission, failed, and cancelled status. A run stopped by abort(), cancelPrompt() or a rejected permission leaves its session cancelled; only a run that failed leaves it failed. Persisted non-terminal activity is repaired from canonical Agent generation outcomes after process restart; incomplete messages and parts normalize to cancelled unless an accepted hidden terminal stage can be promoted.
Conversation Import
A finished conversation from another tool can be imported as an archived, read-only session. The import is staged in three steps, so a large conversation crosses a process boundary in bounded pages. Turning a source format into Flex messages is the importer's job; the core validates, bounds, persists and protects the result.
const started = await harness.beginSessionImport('project:billing', {
sessionId: 'claude-7f3c2a',
title: 'Fix the billing export',
createdAt: '2026-09-01T09:00:00.000Z',
updatedAt: '2026-09-01T11:00:00.000Z',
origin: {
source: 'claude-code-transcript',
sourceId: '7f3c2a',
digest: 'sha256:9b1d…',
details: { cwd: '/workspace' },
},
});
if (started.status === 'started') {
try {
for (const page of pages) {
await harness.appendImportedRuns(started.handle, page);
}
const session = await harness.commitSessionImport(started.handle);
} finally {
await harness.abortSessionImport(started.handle);
}
}
-
beginSessionImport(scopeId, options)validates the options and reserves the caller-chosensessionId. Choose a deterministic id, derived from the source. The result is one of three:started, with a handle{ importId, scopeId, sessionId };already_imported, with the existing session;digest_mismatch, with the existing session.
A scope holds at most one import per
origin.source,origin.sourceIdandorigin.part.index, whatever its session id.already_importedmeans that session has the sameorigin.digest, so the caller skips it.digest_mismatchmeans the source changed since; the caller deletes that session and imports again. A staged import of the same session id or source identity throwsFlexHarnessSessionBusyError. Until the commit completes,listSessions()omits the session andcreateSession()refuses its id. While the import is staged,getSession()throwsFlexHarnessSessionBusyError. While its commit runs,getSession()throwsFlexHarnessNotFoundError, because a deletion tombstone reserves the id. -
appendImportedRuns(handle, runs)stages one page of runs, in order. A run is{ user, assistant }, two messages of the exactIFlexMessageshape. Both belong to the imported session, share onerunIdand one terminal status (completed,failedorcancelled), and have norunningreasoning or tool part. Timestamps are ISO 8601 asDate.prototype.toISOString()renders them. Run ids and message ids are unique within the import. A page that breaks a rule or a bound stages nothing, and the import stays open. The result reports the stagedruns,messagesandbytes. -
commitSessionImport(handle)persists the runs and publishes the session, with onesession.createdevent:originis{ kind: 'import', source, sourceId, digest, part?, details?, importedAt }, andarchivedAtequalsimportedAt;createdAtandupdatedAtare the ones passed tobeginSessionImport(), and the status isidle.
The commit first persists a deletion tombstone for the id, then writes the projection, then replaces the tombstone with the session in one scope save. A failed commit removes what it persisted through the normal tombstone cleanup. After a crash, or a save whose outcome is uncertain, the next load of the namespace resolves it: a surviving tombstone is cleaned up, and a published session is kept. A commit needs at least one run.
-
abortSessionImport(handle)discards a staged import. Nothing of it was persisted. An unknown or finished handle is ignored, so the call is safe in afinally. A staged import that receives noappendImportedRuns()call forstagedImportIdleTimeoutMsexpires and is discarded the same way; its handle then reportsFlexHarnessNotFoundError. -
Child sessions.
parent: { sessionId, runId, toolCallId, agent }imports a child session, such as a subagent's conversation. The parent must be a committed imported session in the same scope. Its assistant message ofrunIdmust hold the tool calltoolCallIdwhosechildSessionIdnames the child. No other session may claim that tool call. The child getsdepthone below its parent, at mostFLEX_SESSION_IMPORT_MAX_DEPTH(8). Deleting the parent deletes its children.
Read-only. An imported session never runs, and nothing writes model context for it: it has no Agent events and no tool jobs. Every operation that would change its history or model context throws FlexHarnessImportedSessionError (code FLEX_SESSION_READ_ONLY):
prompt,startPrompt,enqueuePromptandschedulePrompt;executeSlashCommandandcompactSession;undoSessionandredoSession;pushRuntimeEvent,reconcileToolExecutionandarchiveSessionEvents;- the project-management writes: tasks, goal and scratchpad.
listSlashCommands() reports every command unavailable with the reason Session is an imported conversation and is read-only. Reads work: sessions, messages, message pages, reversion info and project state. So do updateSession() (title and archived flag), deleteSession() and deleteSessionGenerationCohort(). A load never repairs an imported session's runs from Agent generation outcomes, because it has none.
Bounds. sessionImportLimits sets the bounds; each must be an integer from 1 through its maximum:
| Limit | Default | Maximum |
|---|---|---|
maxOpenImports (staged imports at once, per harness) |
2 | 64 |
maxMessagesPerSession |
2048 | 20000 |
maxSessionBytes (JSON bytes of all messages) |
64 MiB | 256 MiB |
maxPageBytes (JSON bytes of one appendImportedRuns() page) |
1 MiB | 1.5 MiB |
stagedImportIdleTimeoutMs (idle time before a staged import expires) |
5 minutes | 1 hour |
Some bounds are fixed:
- one message is at most
FLEX_SESSION_IMPORT_MAX_MESSAGE_BYTES(480 KiB), and must read back through the message transfer limits; origin.detailsis a JSON object of at mostFLEX_SESSION_IMPORT_MAX_DETAILS_BYTES(16 KiB);origin.sourceis at most 128 UTF-8 bytes,origin.sourceId512 andorigin.digest256, none with control characters;origin.parthas acountfrom 1 through 1024 and a zero-basedindexbelow it.
Staged runs are held in memory until the commit, so the memory of open imports is bounded by maxOpenImports × maxSessionBytes. Retiring the scope, disposing the harness or the idle expiry discards staged imports. The expiry timer does not keep the process alive.
Stores. The import adds no store method. Scope snapshots carry the optional session origin, and a store that validates session keys exactly must accept it. The commit writes the whole projection in one save at revision 1. FlexHarness versions before 9.5.0 reject the new key under their strict reader, so a downgrade over imported sessions is unsupported. The tombstone that reserves an import, or that deletes an imported session, carries the optional imported: true, and a store that validates tombstone keys exactly must accept it. FlexHarness 9.5.x rejects a scope snapshot that holds such a tombstone; one exists only while an import commits or an imported session is deleted, or after either was interrupted.
Permissions
Tools request permission through the run-scoped provider context:
await context.requestPermission({
kind: 'filesystem.write',
description: 'Write generated files into the project',
toolCallId,
rememberKey: 'filesystem.write:project-output',
metadata: { target: 'generated/' },
});
Pending requests are runtime-only and queryable with listPendingPermissions(). Every IFlexPermissionRequest carries the exact sessionGenerationId and sessionGenerationSequence of its run so a controller can fence replies against reused session IDs. A controller answers with:
once: allow this request.always: allow and remember the request'srememberKeyfor this session.reject: reject the tool execution.
createFlexToolPermissionRequester(context, { policy }) connects the permission requests of the /tools factories (write, delete, shell, shell-background) to a run: a type the policy allows passes, one it denies fails the tool call with FlexHarnessToolPermissionDeniedError, and every other one becomes a pending request with the tool's title as description and its metadata. Its default rememberKey is the type and a digest of the metadata, so always covers exactly the same request again.
always is invalid when the request has no rememberKey. Remembered decisions are persisted before the waiting tool resolves. If persistence fails, the key is rolled back and the request remains pending so the response can be retried. Concurrent response attempts are serialized and exactly one successful response settles a request.
reject stops the run as a decision, not as a failure. The waiting tool call fails with FlexHarnessPermissionRejectedError, a FlexHarnessAbortError with code FLEX_PERMISSION_REJECTED and the rejected permissionId, and the run ends the way abort() or cancelPrompt() ends it; the model is not asked to recover. The run's completion rejects with that error. run.finished and the finished queue entry have status cancelled and an error whose permissionId names the rejected request. The run's terminal user and assistant messages are cancelled; the assistant message carries rejectedPermissionId next to its error, and so does the session's activity; the session's status is cancelled, as after a stop. The refused tool part fails with the rejection as its error and carries the same rejectedPermissionId. Other permissions still pending in the run are rejected with it, and runs it delegated stop with "The parent run was aborted.". As with a stop, a run rejected after its model produced output stays in future context, where the refused call's result says the permission was rejected. A permission policy deny ends the run the same way.
A bound session resolver may return permissionPolicy(context), an optional per-request function that returns inherit, allow, ask, or deny, synchronously or asynchronously. The cached function should capture its immutable revision; each call receives a fresh frozen context with the exact session generation, run, agent, configurationRef, current run AbortSignal, and a frozen candidate request with a stable permissionId and final resource-namespaced kind and metadata. FlexHarness evaluates it before checking remembered grants, including for built-in subagent.start. inherit follows the existing remembered/pending flow. allow permits this request once without persisting a grant. deny stops the run with a permission rejection, as a reject response does, and creates no pending request. ask remains pending even if a matching grant was remembered; it has requiresExplicitResponse: true and no rememberKey, so always is invalid. An automatic responder must leave such a request for an explicit decision. A controller-side admission needed before allow, such as a delegated-session ticket, must finish before the callback returns, and the host must revoke that admission if the callback's signal aborts or ownership is lost. A cancelled or late callback result cannot approve the run. Callback failures and invalid decisions abort the run through the permissionPolicy external-error source. Immediate policy allow and deny each emit permission.resolved with origin: 'policy', without a preceding permission.requested; deny emits it before stopping the run. The generic callback does not interpret a host's permission grammar.
Tool Call Order
The model often issues several tool calls in one response. FlexHarness runs the read-only ones concurrently and every other call in the order the model issued it: a call that is not read-only starts once every call before it settled, and the calls after it wait for it. The AI SDK alone starts all of a response's calls together once the response ended.
A call is read-only when it has no effect but its result:
- a catalog method declared
readOnly: true, called throughexecuteor from arun_codeprogram (/typedrequestdeclares the methods itsreadOnlyoption marks); - a call the tool handle's
isReadOnlyToolCall({ toolName, input })classifies as read-only, by the tool's name or a catalog method's path and the call's input or params; search_toolsanddescribe_tools.
Every other call, run_code and the built-in tools included, is treated as one that writes. A catalog method's onCall does not override its read-only classification: a read-only declaration covers the hook as well as the method, so the hook may decide concurrent reads and may start while the model still streams. The hook decides every validated call before its result is shared with an identical earlier read. An early execute that proposes waits until the model response confirms that call before handing it to onProposals; a discarded call proposes nothing. Leave a tool or method that asks for permission undeclared too, or the prompts of concurrent calls overlap. Inside a run_code program the same order holds for the calls it makes, in the order it made them; its proposals are still handed over only after the program succeeds.
provideTools: () => ({
tools: { read_file: readFileTool, write_file: writeFileTool },
catalog,
isReadOnlyToolCall: ({ toolName }) => toolName === 'read_file' || toolName.endsWith('.get'),
}),
Two toolbox factories classify their own tools. /tools' isReadOnlyFilesystemToolCall is true for read_file and list_directory of createFilesystemTools(). /mcp's createMcpTools() returns isReadOnlyToolCall, true for a tool its server annotates readOnlyHint: true when the server's configuration sets trustReadOnlyHints: true: an annotation is the server's own claim, so no server's is trusted unless the host says so; createMcpToolCatalog() declares the same tools' methods readOnly. Such a call still passes authorizeToolCall, and when that asks for permission, the requests of concurrent reads are pending together.
provideTools: async () => {
const mcp = await createMcpTools({
servers: { docs: { type: 'streamableHttp', url: docsUrl, trustReadOnlyHints: true } },
});
return {
tools: { ...createFilesystemTools(context), ...mcp.tools },
isReadOnlyToolCall: (call) => isReadOnlyFilesystemToolCall(call) || mcp.isReadOnlyToolCall(call),
close: mcp.close,
};
},
Starting early. A read-only call starts as soon as its input is complete, while the model still streams the rest of its response, when no call that writes precedes it in that response. The response then adopts it: the call keeps running, its tool part starts and its result is recorded as for any call. A response that fails, is retried, or ends without executing the call discards it: its abortSignal fires, and nothing of it reaches the transcript, the events or the model. A permission request of a call the model did not finish issuing waits until the response completed: a catalog method's always, a tool's own when it names the call's toolCallId. A read-only tool sees the confirmation as confirmed in its execution options (IAgentToolExecutionOptions).
Repeated reads. Within a turn, a read-only call identical to an earlier one, by real name and parameters, does not run again: it shares the earlier call's output, and one still running is shared as it finishes. Its tool part carries deduplicatedFrom, the ID of the call that answered it. A call that writes clears what the turn read; a read that failed, or whose response was discarded, is not kept; a tool whose output streams is not shared. A session store that holds such a part is read only by this release and later.
Tool Output Safety
FlexHarness wraps every provided tool execute method before AgentSession receives it. Direct outputs and every AsyncIterable yield are converted into bounded JSON-safe values. Circular references, functions, symbols, bigint values, dates, URLs, and binary values receive deterministic descriptions or records. Returned error objects and unreadable getter values receive fixed descriptions without their original messages. Thrown errors and iterator failures remain failures but are converted to the safe external-error projection before AgentSession observes them.
toolOutputLimits in the complete setup above bounds traversal depth and encoded bytes. The normalizer enforces its byte allowance incrementally: oversized strings are replaced before entering output, and arrays/objects stop reading entries once only truncation metadata fits.
Streaming callbacks use run-local synchronous state rather than one persistence promise per source delta. Text and reasoning accumulate only in the run-local terminal projection while each source delta remains an immediate exact public event. Every distinct async-iterable tool output appears immediately as a bounded cumulative part.updated snapshot while the tool remains running, including the final yielded value before completion. Only the authoritative part.completed output enters the terminal projection, and failed or interrupted tools discard their transient output. callbackLimits bounds callback events, accumulated output bytes, and part count; overflow aborts internally with FlexHarnessCallbackOverflowError and the turn is recorded as failed. The calls of a run_code program never overflow it: they stop their program at its callbacks limit instead. Reservation and terminal finalization are the normal persistence checkpoints, with permission state changes as explicit additional checkpoints.
Model resolver, tool provider, delegated run admission provider, AgentSession, tool execution, tool callback, tool cleanup, and run-persistence failures cross an untrusted error boundary. By default they become a fixed immutable FlexHarnessExternalError before completion rejection, persistence, events, or detached-cleanup reporting. Raw external messages and aggregate members are not retained. A failed onToolCallFinish callback stores and accounts for only the bounded projected message; it does not otherwise reject completion, although exceeding the configured callback limits still fails the run. Scope resolution and the initial store load happen before a run exists and remain outside this boundary.
externalErrorProjector receives one of modelResolver, permissionPolicy, toolProvider, delegatedRunAdmissionProvider, agentSession, toolExecution, toolCallback, toolCleanup, persistence, slashCommand, or turnReversion as its source. It may synchronously return an application-approved plain data object { name, message, code?, limit? }, limited to a 128-byte name, 2048-byte message, optional 128-byte code, and an optional limit that must be a valid IModelLimitInfo. Accessors, extra keys, throwing projectors, and malformed or oversized results fall back to the default projection.
A model call that fails at a provider limit reaches the projector as the typed ModelLimitError with source agentSession; recognize it with isModelLimitError(). Without a projector, or when the projector's result is invalid, FlexHarness projects it itself: code FLEX_USAGE_LIMIT or FLEX_RATE_LIMIT, a message such as Usage limit reached; resets at 2026-09-24T15:00:00.000Z., and the limit. modelLimitErrorInfo(error) returns that projection, so a projector can keep it for limits and redact everything else. The projected limit travels on the rejected FlexHarnessExternalError, the run.finished error, the prompt queue entry error, both failed messages (message.limit) and the session activity (activity.limit), and it survives restarts. Every other external failure still becomes the fixed error. Even exported FlexHarness error subclasses thrown by external integrations are reprojected. Internally created cancellation, callback-overflow, and permission errors retain their typed behavior.
normalizeJsonValue() is also exported for integrations that need the same conversion independently.
The normalized execute result is what the tool part shows. A tool's toModelOutput result is what the model sees: it stays in the canonical private Agent events and is replayed on later steps and turns, also after a reload. It may be an AI SDK v7 content output with file parts, such as a rendered image as { type: 'file', mediaType: 'image/png', data: { type: 'data', data: base64 } }, which OpenAI Responses models receive as input_image. The persisted shapes and their bounds are listed in the Agent runtime readme (@modelprofile.com/flexharness-agent, "Restore a stored session"); a result outside them fails the run.
Agent Runtime Operations
pushRuntimeEvent() appends a validated JSON event through the canonical AgentSession event store. It is intended for controller-owned context such as workspace changes or external notifications; invalid or non-JSON values fail before persistence.
Transactional tool calls persist an execution intent before the tool side effect starts. After an interrupted process, listUncertainToolExecutions() exposes intents whose outcome cannot be proven. A controller must inspect the external system and call reconcileToolExecution() with executed, not-executed, or abandoned-unknown before allowing dependent work to continue. Reconciliation output is normalized using the same tool-output limits.
A model call that fails with a rate limit (HTTP 429), an overloaded provider (529) or an unavailable one (503) is retried at most 8 times, waiting for the provider's retry-after delay but at least the backoff that doubles from 2 s to 30 s. A delay the provider asks for beyond 60 s, or one that would take the call's retries past 150 s in total, fails the run instead: a rate limit (429) fails as a typed rate limit carrying that retry time, a 503 or 529 fails with the provider error. A usage limit is never retried. Each retry emits run.retrying. What the failed call streamed is not kept: its text and reasoning parts leave the assistant message before run.retrying, and the retry streams its own. Tool calls it executed stay, with their results.
agentSessionPolicy forwards bounded AgentSession session controls for context building, compaction, event retention, change-listener pressure, lease cleanup, archived transaction tombstones, and context-overflow retries. A configured contextCompactor receives the projected model messages, only the filtered model-visible covered events, AgentSession's existing reason, abortSignal and reportUsage, and the exact resolved scopeId, scope, storageKey, and sessionId for the invocation causing compaction. The invocation context remains isolated when aliases share one storage key, so integrations can resolve the correct model without global mutable state. If no events are eligible for compaction, compactSession() returns without calling the compactor or writing a compaction event; otherwise it writes the canonical event. archiveSessionEvents() moves events covered by that compaction into the configured Agent event archive store and returns public archive metadata.
With a compactor, proactive compaction defaults to a 256 KiB UTF-8 JSON model-context budget before inference. agentSessionPolicy.contextCompactionBytes changes this byte budget (1 through 32 MiB); false disables it. The reason is context-budget, and resolvedModel names the current run's model, including delegated runs. Only settled history is summarized; the current prompt and unsettled generation remain intact. Once the generation is finalized and no uncertain tool intent remains, automatically compacted source events move into the durable archive when the store supports it. Visible transcript messages remain available. Manual compaction remains explicitly archiveable and does not automatically move its source events.
executionContextProvider can construct a AgentSession execution context for each session. FlexHarness supplies the resolved scope, storage key, and the session's private job store. The public background APIs expose only execution ID, type, state, exit code, and timestamps; command payloads, stdout, and stderr remain private. The provider's optional close() is owned by session deletion, scope retirement, and harness disposal.
Context Compaction Model
A contextCompactor states which model to compact with from the model the compacted work was dispatched on, never from a host default that may name another account:
reason |
Compaction | Model members |
|---|---|---|
context-overflow |
A run's context overflowed, including a delegated run's | resolvedModel: the IFlexResolvedModel the model resolver returned for that run, frozen; a delegated run's overflow carries the delegated run's own model |
context-budget |
Settled history exceeds contextCompactionBytes before a generation's first inference (256 KiB by default with a compactor) |
resolvedModel: the exact model resolved for that run, including its delegated model |
manual |
/compact or compactSession() |
modelHint: the modelHint of the /compact execution options or of compactSession(scopeId, sessionId, { modelHint }), when one was passed |
retention |
Event retention | none |
A compactor uses resolvedModel.model directly, and resolves modelHint the way its model resolver resolves a prompt's modelHint. Only when neither member is present, because nothing dispatched states a model, does the host choose the compaction model itself.
A context-overflow or context-budget compaction also carries system and providerOptions: the instructions and the provider options (cache options applied, so store: false on OpenAI) the run's turns are sent with. A context-budget compaction on a provider that caches the prefix also carries the turns' tools. compactMessages() sends its summary call with them as one streamed call, the shape of every turn, with the tools' definitions but none of their code, so a compactor that passes its options through repeats the run's prefix and reads it from the provider's prompt cache; definitions that outweigh the history nine times are left out. Without system it sends its own instructions. With an attachmentResolver on the harness, every compaction also carries attachmentResolver, bound to the session, so compactMessages() sends the history's referenced files as the turns do.
agentSessionPolicy: {
contextCompactor: async (messages, _events, options) => {
const model = options.resolvedModel?.model
?? await resolveHostModel(options.scopeId, options.modelHint);
return compactMessages(model, messages, options);
},
},
Events
const unsubscribe = harness.subscribe((event) => {
switch (event.type) {
case 'part.delta':
applyExactDelta(
event.sessionId,
event.messageIndex,
event.partIndex,
event.partType,
event.delta,
event.baseTextUtf8Bytes,
event.textUtf8Bytes,
);
break;
case 'part.removed':
removePart(event.sessionId, event.messageIndex, event.partIndex);
break;
case 'permission.requested':
showPermission(event.request);
break;
case 'session.updated':
renderSession(event.session);
break;
case 'session.deleted':
removeSession(event.sessionId);
break;
case 'run.retrying':
showRetry(event.runId, event.reason, event.delayMs, event.attempt, event.maxAttempts);
break;
case 'model.usage':
if (event.status === 'reported') countUsage(event.scopeId, event.provider, event.requestedModelId, event.usage);
else markUsageIncomplete(event.scopeId, event.runId, event.reason);
break;
case 'run.finished':
markRunFinished(event.runId, event.status, event.error?.limit);
// A run that did not stay in context needs its prompt sent again, followed by its unapplied steers.
resendAsNextTurn(event.keptInContext ? [] : [event.runId], event.unappliedSteerIds);
break;
case 'prompt.steer.accepted':
markSteerPending(event.steerId);
break;
case 'prompt.steer.applied':
markSteerApplied(event.steerId, event.messageId);
break;
case 'session.history.changed':
refreshSelectedSession(event.sessionId);
break;
case 'prompt.queued':
case 'prompt.started':
case 'prompt.running':
case 'prompt.finished':
renderQueueStatus(event.queueId, event.entry.status);
break;
}
});
unsubscribe();
Events are discriminated, deeply immutable values with one global sequence within each FlexHarness instance. Every IFlexEventBase carries the exact source sessionGenerationId and sessionGenerationSequence; session.deleted retains the deleted generation after live metadata is removed. Listener exceptions are isolated from runs and other listeners. Every accepted queue entry emits prompt.queued and exactly one prompt.finished. Every accepted steer emits prompt.steer.accepted, and prompt.steer.applied, after the message.* events of its user message, once the run took it in; run.finished lists it in unappliedSteerIds unless it is in the model context of later turns. Durable promotion additionally emits prompt.started, and actual model preparation emits prompt.running; cancellation or failure can omit either intermediate event. Existing durable run/message terminal events precede prompt.finished. Every callback-backed streamed text part that stays in its message emits exactly one part.completed event before the corresponding run.finished event: once the model step that streamed it is committed, so its text is final and saved, or with the run's terminal projection when that step did not commit. Text a later model step streams starts a new part. When a model call fails and is retried, the text and reasoning parts it streamed are not in the session: part.removed takes each of them out of its message, last part first, and the retry streams its own parts. Its partIndex is the part's position before it left; the parts after it move up by one. A removed part has no further events, and part.removed does not count against callbackLimits. The tool parts of calls the failed model call executed stay. run.retrying follows those removals and precedes each wait before a rate-limited, overloaded or unavailable model call is retried. It carries the one-based attempt, maxAttempts, delayMs and reason (rate_limit, overloaded or unavailable) and counts against callbackLimits. Every part.started, part.delta, part.updated, part.completed, and part.removed event carries zero-based messageIndex and partIndex coordinates from the session's authoritative message and part sequences.
getSessionView(scopeId, sessionId) reads a session's metadata, visible messages and pending permission requests in one step, together with the instance's event sequence they include. The messages include what a running turn has streamed so far, so a consumer that applies every event of the instance with a greater sequence follows the session exactly from the view on; events up to sequence are already in it. The read waits while a queued write or a change that emits events for the session is in progress; a session that keeps changing through 64 consecutive waits is refused with FlexHarnessSessionBusyError, and the read can be repeated. Session metadata is current: status changes such as waiting_permission emit no event. A permission request its run leaves unanswered leaves pendingPermissions without an event of its own; the run's run.finished follows.
model.usage reports each model call of a session once, whatever the outcome of its run: status: 'reported' with the provider's usage, responseModelId and cacheHitRatio (the share of the call's input read from the provider's prompt cache), or status: 'unreported' with reason aborted, failed or missing for a call that ended before its provider reported usage; the provider may still have consumed tokens for such a call. Every event carries provider and requestedModelId, the model the call was made with; key usage caps on those two. source: 'generation' is a model step of the run runId, retried calls included. source: 'compaction' is a model call of a context compaction that the contextCompactor reported through the reportUsage in its options; runId names the run whose context overflow caused it and is absent for compactSession() and event retention. compactMessages(model, messages, options) from @modelprofile.com/flexharness/compaction reports each attempt of its call when the compactor passes its options through. A compactor's model calls that it does not report, and model calls a tool makes itself, are not counted. model.usage events do not count against callbackLimits. A run finishes only after its generation has ended, and a generation ends only after it has reported each of its model calls, also when the run fails or is cancelled. So every model.usage event of a run precedes its run.finished, and the run's terminal usage includes every reported call. For a compaction call this holds when the contextCompactor reports it before its promise settles. This wait is bounded: a model call still open MODEL_ABORT_GRACE_MS (1 s) after the run's abort, because its transport ignores the abort signal, fails with the signal's reason, is reported unreported with reason aborted, and its stream is cancelled; a transport that heeds the abort and still sends its usage within that time is counted. The terminal assistant message of a completed, failed or cancelled run carries usage, the sum of the run's reported calls; for a failed or cancelled run it is a lower bound when one of the run's model.usage events is unreported.
part.started, part.updated, and part.completed are snapshot events with a complete immutable part. Tool and reasoning parts carry startedAt (ISO time the call or reasoning began) and, once they reach a terminal status, completedAt; parts recorded by earlier versions have neither, so a consumer falls back to the message times for them. A newly streamed text part emits part.started with empty text before its first delta; reasoning parts also start empty. part.delta is a separate delta-only event: it has no cumulative part, and carries partType, the required exact source delta, and baseTextUtf8Bytes/textUtf8Bytes for the cumulative text before and after that delta. The counters remain correct when a UTF-16 surrogate pair is split across callbacks. Exact deltas are never truncated, including a single delta above the 96 KiB message-transfer text limit. part.updated remains a cumulative replacement snapshot for running tool output or metadata, not a text delta. session.history.changed carries direction: 'undo' | 'redo' | 'branch'; runId is present for undo and redo. Events contain public IDs, exact delta payloads, and public snapshots only; they do not expose prompt payloads, the resolved scope object, storage key, model object, provider options, or raw storage key.
Part events narrow through the exported TFlexPartEvent union. IFlexPartEventBase contains their shared coordinates, IFlexPartSnapshotEvent owns snapshot events and their complete part, IFlexPartDeltaEvent owns exact delta-only events and their UTF-8 counters, and IFlexPartRemovedEvent is part.removed, which carries only the coordinates.
Migrating Part Events to 5.x
Version 5.x replaces cumulative part.delta payloads with the exact delta-only contract above. Consumers must stop reading event.part from part.delta; use event.partType, event.delta, event.baseTextUtf8Bytes, and event.textUtf8Bytes, then hydrate or settle from the complete part carried by snapshot events. IFlexPartChangedEvent has been removed; use TFlexPartEvent, IFlexPartSnapshotEvent, or IFlexPartDeltaEvent according to the required narrowing. New streamed text and reasoning parts start with empty text, so consumers must apply subsequent deltas in sequence within that harness instance.
Stores
Current FlexHarness persistence is separated by trust and lifecycle domain through IFlexHarnessStores:
scopes: session metadata and deletion tombstones for a resolved storage namespace.projections: public audit messages and hidden terminal stages per session.permissions: remembered permission keys per session.projectManagement: generation-fenced task, goal, and scratchpad state per session.agentEvents: canonical private AgentSession events and archives per session.jobs: private background execution state per session.results(optional): tool outputs over the output budget, per session, written once under their handle and listed oldest first, so FlexHarness can delete all but a session's latest (Large results). Without it such outputs are truncated.
InMemoryFlexHarnessStores implements all six required domains and results with revision-based compare-and-swap behavior for tests and ephemeral processes. It is the default when stores is omitted. Custom IFlexHarnessStores implementations must provide projectManagement even when project-management tools are disabled, because deletion cleanup always writes the generation fence.
A custom store validates what it loads and saves with the validators the built-in stores apply. Each throws FlexHarnessStoreFormatError for a malformed or wrong-schema snapshot: assertFlexScopeSnapshot(), assertFlexProjectionSnapshot(value, sessionId), which also requires every message and staged terminal to belong to sessionId, assertFlexPermissionSnapshot(), assertFlexProjectManagementRecord() and assertFlexToolJobSnapshot(). Agent event snapshots and archives have validateAgentEventSnapshot() and validateAgentEventArchive() in @modelprofile.com/flexharness-agent.
Projections and Agent event histories are saved as changes, so a save costs what changed rather than the whole session. IFlexProjectionStore.save(storageKey, sessionId, change, expectedRevision) receives an IFlexProjectionChange: the schema-3 projection with messages and reversionSegments given as IFlexListChanges { start, end, items }, each of which keeps the stored entries from start up to end and appends items. A list change with start and end at 0 replaces the list whole; FlexHarness sends whole lists when it cannot know which entries the store holds: for a projection it loaded in an older schema and after a save with an uncertain outcome. The Agent event store of each session receives an IAgentEventChange of removed, replaced and appended events, as the @modelprofile.com/flexharness-agent readme describes. A store that keeps snapshots checks a change's form with assertFlexProjectionChange() or assertAgentEventChange(), applies it with applyFlexProjectionChange() or applyAgentEventChange(), and validates the result as above; createWholeProjectionChange(snapshot) makes the change that stores a projection whole. FlexHarness freezes every message and reversion segment it saved and replaces an entry instead of editing it, which is how it tells the entries the store already holds. /stores/testing exports createFlexStoreContractTests(backend), the cases every built-in store passes, for a custom store to run with any test runner; its readme shows how.
Custom Agent event and job providers may implement releaseSession(storageKey, sessionId) to release session-bound wrappers, handles, or caches without deleting durable data. FlexHarness calls these hooks only after the corresponding AgentSession or execution context has released runtime ownership. A failed release remains owned for a later retirement or disposal retry. deleteSession() and deleteSessionGenerationCohort() are the separate destructive operations for durable session data.
JsonFileFlexHarnessStores stores the domains in separate scopes, projections, permissions, projectManagement, events, archives, jobs, and results directories. Storage and session identifiers are SHA-256 hashed for filenames. Passing the store bundle supplies lifecycle persistence but does not enable any built-in tool. It provides:
- Strict domain-specific schema validation and optimistic revisions.
- Static process-wide queues shared by all store instances for the same absolute file.
- Revision re-reads inside the queue before every save.
- Atomic temporary-file write, file fsync, rename, and parent-directory fsync.
- Directory mode
0700and file mode0600, including existing paths. - Stale temporary-file cleanup and strict snapshot validation.
After every harness using a JsonFileFlexHarnessStores instance has been disposed and no store operation remains active, call await stores.dispose() to retry and drain any file handle whose earlier close failed. A failed store disposal retains that handle so the call can be retried.
The JSON stores are explicitly not cross-process safe. When several processes can access the same storage namespace, every core IFlexHarnessStores domain and stores.projectManagement must use database-backed or equivalent cross-process CAS. Process-local CAS for the core stores or for PM alone is insufficient: session generation creation, cleanup tombstones, PM replacement, and stale-writer rejection must all retain their respective atomic preconditions across processes.
NoSqlFlexHarnessStores from /stores/nosqldb is that database-backed bundle for every domain. Pass it an initialized SmartdataDb and await stores.prepare() at startup, before the first harness uses it. Its purgeNamespace(storageKey) removes a retired scope's records in every domain, and purgeNamespace(storageKey, { fence: true }) also refuses every later record of that namespace, so a harness that still holds the scope cannot write it back; the subpath readme describes chunked snapshots, change logs, namespace purges and failure reporting.
Direct store operations and non-run session mutations surface conflicts as FlexHarnessStoreConflictError or the corresponding AgentSession store conflict. Malformed, wrong-schema, or non-JSON snapshots are surfaced as FlexHarnessStoreFormatError. A write or deletion that changed its target but cannot confirm parent-directory durability surfaces FlexHarnessStoreCommitUncertainError with the affected path, operation, and cause. A record creation in a namespace that a fenced NoSQLDB purge removed surfaces FlexHarnessStorePurgedError with the refused path. Run persistence failures cross the external error boundary and therefore become FlexHarnessExternalError. FlexHarness does not merge conflicts.
Each domain serializes its own mutations. A successful run first persists a hidden completed projection, then finalizes the canonical Agent generation as accepted, then promotes the hidden projection publicly. Recovery uses the canonical generation outcome to promote an accepted stage or publish a failed/cancelled projection. Failed and cancelled generations remain auditable but never enter future model context.
Projection stores accept schema version 1, 2, or 3 from load(), but every save() receives a change that makes the current schema-3 shape. A change to a loaded schema-1 or schema-2 projection carries both lists whole. Schema 3 records explicit reversion protocol, transcript/workspace provenance, and V2 disposition. A terminal V2 segment must be conclusively revertible, no-change, or nonrevertible; a pending V2 capture remains owned by its capture WAL and is never treated as transcript history.
Current schema-3 projections also persist optional metrics state. Older strict readers that do not recognize this key cannot reopen a projection after metrics have been written; downgrade or mixed-version access to that session is unsupported.
A loaded schema-1 projection has no reversion state. Schema-2 workspaceCaptured segments migrate as protocol-1 workspace/revertible history without discarding references; transcript-only segments migrate as protocol-1 transcript provenance. Under workspace-required, that legacy transcript history is a barrier. The next projection mutation writes schema 3. Custom stores must preserve strict compare-and-swap revisions across schema-1 and schema-2 read-upgrade-write cycles, including uncertain-save reconciliation.
Migrating From 2.x
Version 3.x replaces the single IFlexHarnessStore snapshot with the split stores above. Run migration while every process that can access the storage namespace is stopped.
import {
JsonFileFlexHarnessStores,
} from '@modelprofile.com/flexharness';
import {
migrateLegacyFlexHarnessSnapshot,
type IFlexLegacyHarnessSnapshot,
} from '@modelprofile.com/flexharness/migration';
const storageKey = 'account/project';
const legacySnapshot: IFlexLegacyHarnessSnapshot = await loadLegacySnapshot(storageKey);
const stores = new JsonFileFlexHarnessStores({
directory: '/var/lib/my-app/model-sessions-v3',
});
await migrateLegacyFlexHarnessSnapshot(storageKey, legacySnapshot, stores);
loadLegacySnapshot() is application-owned access to the snapshot written by the 2.x store. The migration validates the complete source and every public run before writing. A run left streaming by a process crash is deterministically repaired to the same cancelled state that the 2.x loader produced in memory. The migration then converts private model messages into generationless canonical Agent conversation events, records terminal AgentSession transactions for completed, failed, and cancelled public runs, and writes schema-3 projections with empty reversion state. Migrated history therefore remains visible and auditable but is not retroactively undoable; turns created after migration receive normal reversion segments and optional workspace captures. The migration preflights the scope, projection, permission, Agent event, and job destinations before any write, applies missing per-session domains first, and publishes scope discovery last. It is safe to rerun after no work, a completed prefix, or a complete migration when existing destination content is identical. It fails closed when a destination contains conflicting content or non-empty jobs. Keep the legacy snapshot until the migrated application has loaded and verified every storage namespace.
Shutdown
Model and tool resolution share the run signal. Synchronous throws are observed as resolver failures; the first failure aborts that signal and finalizes immediately without waiting for an unresponsive sibling. A detached application or resource tool provider that resolves later is observed and every acquired handle is closed; disposal waits for that settlement and reports a late close failure.
Tool-handle close settles before a turn can be successful. After model generation resolves, FlexHarness stages its terminal projection before canonical acceptance; earlier execution failures are finalized as interrupted and then published from that durable outcome. Cleanup failure prevents canonical acceptance. If canonical finalization or public promotion fails, FlexHarness fences the namespace; the next load repairs public state from the durable canonical outcome and any hidden terminal stage. The run still settles in the process that ran it: its completion rejects, and listeners receive its terminal message.updated events and one run.finished with status failed (cancelled for an owner cancellation or a rejected permission) before prompt.finished. Those events are not persisted; a reload shows the repaired state, where a run without a durable outcome appears interrupted. keptInContext on such a run.finished is true only when the run's accepted outcome is already durable.
dispose() is asynchronous and idempotent. It marks the harness closed, settles unresolved queue admissions, cancels queued prompts and cancellable active runs, rejects pending permissions, emits queue terminal events, waits for committing runs, queue drains, all run finalizers, state save tails, and tracked detached tool-provider cleanup, purges runtime queue records, then clears listeners. Loaded state caches are cleared after all cleanup succeeds; a failed drain retains its cache and cleanup ownership so a later dispose() call can retry it. Multiple run or cleanup failures are reported through FlexHarnessRunError.
If dispose() overlaps a storage namespace already being retired, both calls await the same storage drain and cleanup runs once. A retirement call begun after disposal starts rejects with FlexHarnessClosedError.
Cancellation is cooperative: model resolvers, tool providers, AgentSession model execution, tools, execution contexts, and cleanup functions must observe the supplied AbortSignal and settle tracked work. A model call is the exception: one whose transport ignores the signal is cut off MODEL_ABORT_GRACE_MS after the abort. After a sibling resolver fails, FlexHarness deliberately does not wait for an unresponsive model resolver; a detached tool provider remains tracked because any late handle must be closed. A process host that needs a hard shutdown deadline must enforce that deadline outside FlexHarness and terminate only the process it owns.
License and Legal Information
This repository contains open-source code licensed under the MIT License. A copy of the license can be found in the repository license file.
Please note: The MIT License does not grant permission to use the trade names, trademarks, service marks, or product names of the project, except as required for reasonable and customary use in describing the origin of the work and reproducing the content of the NOTICE file.
Trademarks
This project is owned and maintained by Task Venture Capital GmbH. The names and logos associated with Task Venture Capital GmbH and any related products or services are trademarks of Task Venture Capital GmbH or third parties, and are not included within the scope of the MIT license granted herein.
Use of these trademarks must comply with Task Venture Capital GmbH's Trademark Guidelines or the guidelines of the respective third-party owners, and any usage must be approved in writing. Third-party trademarks used herein are the property of their respective owners and used only in a descriptive manner, e.g. for an implementation of an API or similar.
Company Information
Task Venture Capital GmbH
Registered at District Court Bremen HRB 35230 HB, Germany
For any legal inquiries or further information, please contact us via email at hello@task.vc.
By using this repository, you acknowledge that you have read this section, agree to comply with its terms, and understand that the licensing of the code does not imply endorsement by Task Venture Capital GmbH of any derivative works.