14 KiB
Harness Auth Preflight + Runtime Failure Detection — Tech Spec
Context
See PRODUCT.md for user-facing behavior. This document covers the implementation.
Current state
Third-party harnesses implement the ThirdPartyHarness trait (app/src/ai/agent_sdk/driver/harness/mod.rs:122). The lifecycle for a third-party harness run is:
ThirdPartyHarness::validate()— checks CLI is on PATH.ThirdPartyHarness::build_runner()— writes config files (auth, trust, system prompt, MCP) and returns aHarnessRunner.HarnessRunner::start()— creates the external conversation on the server and launches the main CLI command.
These steps are orchestrated in AgentDriver::run_internal (driver.rs:1608), which calls setup_harness → prepare_harness → run_preflight_checks → run_harness.
Driver errors are classified in error_classification.rs into (AgentTaskState, TaskStatusUpdate) pairs and reported to the server via report_driver_error. The PlatformErrorCode enum already has AUTHENTICATION_REQUIRED.
Relevant files
app/src/ai/agent_sdk/driver/harness/mod.rs—ThirdPartyHarnesstrait,HarnessRunnertrait,preflight_commands_forapp/src/ai/agent_sdk/driver/harness/claude_code.rs,harness/codex.rs,harness/gemini.rs— per-harness implementationsapp/src/ai/agent_sdk/driver.rs—AgentDriver,AgentDriverErrorenum,run_internal,prepare_harness,run_harness,run_preflight_checksapp/src/ai/agent_sdk/driver/harness_output_monitor.rs— runtime scanner (new)app/src/ai/agent_sdk/driver/error_classification.rs— mapsAgentDriverError→(AgentTaskState, TaskStatusUpdate)app/src/ai/agent_sdk/driver/terminal.rs—TerminalDriver::execute_command,block_snapshot,find_first_match_in_block_output(new)app/src/terminal/model/find.rs—RegexDFAsmachinery reused by the scannerapp/src/terminal/model/blockgrid.rs—BlockGrid::find(&RegexDFAs)returns matches over cell storage directly
Proposed changes
1. Trait surface changes on ThirdPartyHarness
Drop the billing preflight method and add a runtime-error-patterns method:
// Removed:
// fn billing_check_command(&self) -> Option<String> { None }
// Added:
/// Substrings to scan for in the running harness block's output. A hit
/// indicates the harness can't make a successful API request.
fn runtime_error_patterns(&self) -> &'static [&'static str] {
&[]
}
Per-harness:
- Claude Code and Codex: drop
billing_check_command, addruntime_error_patternsreturning an empty slice with aTODO(REMOTE-1385)to fill in validated needles. - Gemini: inherits default
None/ empty.
preflight_commands_for(harness) in harness/mod.rs now only pushes auth_check_command(). The viewer's "setup commands" grouping continues to recognize preflight blocks by exact string equality.
2. New error variant + classification
In driver.rs's AgentDriverError:
// Removed: HarnessAuthFailureKind enum and the `kind` field on
// HarnessAuthCheckFailed.
#[error("Harness '{harness}' auth preflight failed")]
HarnessAuthCheckFailed {
harness: String,
detail: String,
},
// Added:
#[error("Harness '{harness}' reported a runtime failure matching '{pattern}'")]
HarnessRuntimeFailureDetected {
harness: String,
pattern: String,
excerpt: String,
},
In error_classification.rs, both variants map to (AgentTaskState::Failed, PlatformErrorCode::AuthenticationRequired) with distinct user-visible messages. The runtime failure message surfaces both the matched needle and the excerpt from the harness output.
3. New find_first_match_in_block_output helper on TerminalDriver
Add a sibling to block_snapshot in driver/terminal.rs that runs the existing DFA machinery against the block's output grid directly:
pub struct BlockOutputMatch {
pub matched_text: String,
pub excerpt: String,
}
pub fn find_first_match_in_block_output(
&self,
block_id: &BlockId,
dfas: &RegexDFAs,
ctx: &AppContext,
) -> Option<BlockOutputMatch>;
Internally:
- Acquire the terminal model lock (same scope as
block_snapshot). - Look up the block by ID; return
Noneif missing. - Run
block.output_grid().find(dfas).next()to get the firstMatch = RangeInclusive<Point>. - Use the grid handler's
bounds_to_stringto extract the matched substring (used to identify the originating needle) and the full row(s) the match touches (used as the user-visible excerpt). Both withinclude_esc_sequences=falseandRespectObfuscatedSecrets::Yesso the failure message never leaks credentials.
No new grid-side code is required — BlockGrid::find is the same path the find feature uses on every keystroke.
4. New harness_output_monitor module
File: app/src/ai/agent_sdk/driver/harness_output_monitor.rs.
pub(crate) struct DetectedHarnessError {
pub pattern: String,
pub excerpt: String,
}
const SCAN_SCHEDULE: &[(Duration, Duration)] = &[
(Duration::from_secs(5), Duration::from_secs(30)),
(Duration::from_secs(15), Duration::from_secs(60)),
];
pub(crate) fn build_dfas(patterns: &[&'static str]) -> Option<RegexDFAs>;
pub(crate) fn pattern_for_match(matched_text: &str, patterns: &[&'static str]) -> Option<&'static str>;
pub(crate) async fn watch_block_for_errors(
block_id: BlockId,
patterns: &'static [&'static str],
foreground: &ModelSpawner<AgentDriver>,
) -> Option<DetectedHarnessError>;
Implementation notes:
- Resolve immediately with
Nonewhenpatterns.is_empty()so Gemini/Oz pay zero runtime cost. - Build the combined DFA once at scanner entry via
RegexDFAs::new_many(with regex-escaped needles + case-insensitive matching). Wrap inArcso each tick can clone cheaply. - Use
warpui::r#async::Timer::afterfor ticks, matching the existingrun_harnessperiodic-save pattern. - Each tick spawns onto
foregroundto callTerminalDriver::find_first_match_in_block_output(&dfas). OnSome, the scanner runs the stall confirmation loop (see below) before mappingmatched_text→ originating needle viapattern_for_match, capping the excerpt to ~240 chars, and resolving the future withDetectedHarnessError. - The DFA scans cell storage from the start each tick, so no
last_scanned_lencursor is needed — the regex_automata cache memoizes transitions.
Stall confirmation loop
After a pattern hit, the scanner runs confirm_stall before reporting a failure. This guards against false positives when a harness prints a transient API error and then automatically retries.
/// Gap between consecutive plaintext snapshots while confirming a hit.
const STALL_POLL_INTERVAL: Duration = Duration::from_secs(10);
/// Total budget for the stall-confirmation loop.
const STALL_CONFIRMATION_BUDGET: Duration = Duration::from_secs(60);
/// Pure equality check kept out of the async body for easy unit tests.
fn outputs_stalled(before: Option<&str>, after: Option<&str>) -> bool {
matches!((before, after), (Some(a), Some(b)) if a == b)
}
/// Polls the block plaintext every `STALL_POLL_INTERVAL` for up to
/// `STALL_CONFIRMATION_BUDGET`. Returns `(Some(hit), elapsed)` once two
/// consecutive snapshots are byte-identical AND a pattern is still
/// present after the second snapshot; otherwise `(None, elapsed)`.
async fn confirm_stall(
block_id: &BlockId,
dfas: &Arc<RegexDFAs>,
foreground: &ModelSpawner<AgentDriver>,
) -> (Option<BlockOutputMatch>, Duration);
Design notes:
- Why plaintext equality over
dirty_cells_range?GridHandler::dirty_cells_rangeis per-pty-pass-scoped — it resets to the cursor point on everyon_finish_byte_processingcall, so it's not a usable "has anything happened lately" oracle. Sampling the visible plaintext twice and comparing them gives us the same information directly, and works for spinner-only retries (the spinner cell changes between samples → the snapshots differ → the loop keeps running). - Why count confirmation time against the outer scan-schedule budget? A flaky harness that keeps retripping the same pattern would otherwise extend the watch window indefinitely — each candidate hit would buy another 60 seconds. Adding
confirmation_elapsedback to the outerelapsedcounter caps the total observation period at the outer 90-second budget plus whatever final confirmation was in flight when the schedule exhausted. - A new helper
TerminalDriver::block_output_plaintext(block_id, ctx) -> Option<String>is added alongsidefind_first_match_in_block_output. It readsblock.output_grid().contents_to_string(false, None)under the same model lock and obfuscates secrets identically.
5. Wire the scanner into run_harness
run_harness gains two new parameters and a fourth select arm:
async fn run_harness(
runner: Arc<dyn HarnessRunner>,
harness_name: String,
runtime_error_patterns: &'static [&'static str],
foreground: &ModelSpawner<Self>,
harness_exit_rx: oneshot::Receiver<()>,
) -> Result<(), AgentDriverError> {
let command_handle = runner.start(foreground).await?;
let block_id = command_handle.block_id().clone();
let mut command_handle = command_handle.fuse();
let mut harness_exit_rx = harness_exit_rx.fuse();
let scanner_fut = harness_output_monitor::watch_block_for_errors(
block_id, runtime_error_patterns, foreground,
)
.fuse();
futures::pin_mut!(scanner_fut);
let mut detected_runtime_failure: Option<harness_output_monitor::DetectedHarnessError> = None;
let command_result = loop {
futures::select! {
exit_code = command_handle => break exit_code,
_ = warpui::r#async::Timer::after(HARNESS_SAVE_INTERVAL).fuse() => {
/* existing periodic save */
}
_ = harness_exit_rx => {
/* existing graceful-exit branch */
}
detected = scanner_fut => {
if let Some(error) = detected {
// Ask the harness to exit gracefully so cleanup runs;
// record the failure and let the command future complete.
let _ = runner.exit(foreground).await;
detected_runtime_failure = Some(error);
}
// When the schedule exhausts without a hit, the `Fuse`
// wrapper makes this branch stay Pending forever.
}
}
};
/* existing final save + cleanup */
if let Some(error) = detected_runtime_failure {
return Err(AgentDriverError::HarnessRuntimeFailureDetected {
harness: harness_name,
pattern: error.pattern,
excerpt: error.excerpt,
});
}
/* existing exit-code mapping */
}
Notes:
- After
scanner_futresolves toNone, theFusewrapper makes the branch stayPendingforever, so we don't busy-loop. - A detected runtime failure takes precedence over the harness's own exit code: surface the actionable detail rather than a generic "exit code N".
- The cleanup-disposition computation also factors
detected_runtime_failureso we always drop resumption state on a runtime failure (a failed run shouldn't be silently resumable).
6. Caller update in run_internal
The third-party branch in run_internal already has harness.as_ref() in hand. Plumb the harness name and runtime patterns through both the credentials-refresh branch and the no-refresh branch:
let harness_name = harness.cli_agent().command_prefix().to_owned();
let runtime_error_patterns = harness.runtime_error_patterns();
Self::run_harness(runner, harness_name, runtime_error_patterns, &foreground, harness_exit_rx).await
7. Server-side changes
None. Both failure modes map to the existing AUTHENTICATION_REQUIRED PlatformErrorCode, differentiated only by the human-readable statusMessage. The server already handles AUTHENTICATION_REQUIRED in task state transitions.
Testing and validation
Unit tests
harness_output_monitor_tests.rs(new):build_dfasreturnsNonefor empty patterns andSomeotherwise; regex metacharacters in needles are escaped.pattern_for_matchmaps a case-different matched text back to the originating'staticneedle; returnsNonewhen no needle matches; picks the first matching needle when multiple lowercase-equal candidates exist.
error_classification_tests.rs:HarnessAuthCheckFailedmaps to(Failed, AuthenticationRequired)with the auth-failure message. (Existingharness_auth_test_request_failed_*test removed since theTestRequestFailedvariant is gone.)HarnessRuntimeFailureDetectedmaps to(Failed, AuthenticationRequired); the user-visible message contains both the matchedpatternand theexcerpt.
harness/mod_tests.rs:- Billing tests removed (
*_billing_check_command, billing assertions inpreflight_commands_for_*). preflight_commands_for_claude_returns_auth_only/preflight_commands_for_codex_returns_auth_onlypin the auth check string to the trait impl.runtime_error_patterns()is callable on Claude/Codex (slice may be empty); Gemini returns an empty slice.
- Billing tests removed (
Manual validation
- Claude Code with a valid API key → auth check passes, harness runs, scanner finds no match in 90 seconds → run completes normally.
- Claude Code with an invalid API key → auth check fails, task is marked FAILED with the auth-failure message.
- Claude Code with a valid key but exhausted credits → auth check passes, harness emits a credit-balance error to its block, scanner detects it, task is marked FAILED with the runtime-failure message including the matched pattern + excerpt.
- Same matrix for Codex.
- Gemini → auth check skipped, scanner is a no-op, harness proceeds and behaves as before.
Parallelization
Single-agent task. The change is tightly coupled (trait method, classifier, driver loop, scanner module, new helper, tests) and lives in one crate; parallel agents would step on each other.