The previous implementation only used the standard OpenAI /models
endpoint which often lacks context window and capability metadata,
causing a blind 200K fallback for all models.
Now fetch_openai_models_from_endpoint tries LiteLLM's /model/info
endpoint first, which returns rich metadata:
- max_input_tokens (e.g. 1,000,000 for Sonnet 4.6)
- max_output_tokens (e.g. 128,000 for max models)
- supports_vision
- supports_function_calling
- underlying model path (for provider detection)
Falls back to /models if /model/info is unavailable (e.g. non-LiteLLM
OpenAI-compatible endpoints).
This ensures the model picker and context window configuration reflect
the actual capabilities of the configured models.
Galaxy should never call Warp's cloud API for AI. This is a policy
requirement. All model availability is now determined exclusively by
locally configured providers (Bedrock and/or OpenAI/LiteLLM).
Changes:
- Disable refresh_authed_models, refresh_public_models, refresh_available_models
(now no-ops with debug log)
- Disable on_server_update and update_feature_model_choices
- Disable get_cached_models (no stale server models restored from cache)
- Replace ModelsByFeature::default() with minimal placeholder that gets
stripped by inject_bedrock_models/inject_openai_models
- Make inject_bedrock_models strip Unknown placeholders unconditionally
- Make default_llm_info() return a static fallback instead of panicking
when no models are configured (prevents null reference crashes)
- Add has_any_provider_models() for UI to check provider availability
- Make ProviderConfig::None return a user-friendly error instead of
calling Warp's cloud API (the previous fallback behavior)
Safety: if no providers are enabled, the system gracefully returns an
error message rather than crashing or silently calling Warp's servers.
- Update version from 1.6.3 to 2.0.0 in app/Cargo.toml and Cargo.lock
- Add install-galaxy.sh upload step to build-and-deploy-hermes script
- Include pending AI provider and agent changes
Remove duplicate popup rendering from render_vertical_tabs_panel.
The popup was rendered both inside the panel's stack AND at the
workspace level in a Dismiss overlay, causing event dispatch conflicts
due to shared MouseStateHandle instances between the two identical
popup trees.
Major changes:
- **Bedrock translator architecture**: Extract orchestration logic from `impl.rs` into
a dedicated `translator.rs` module. Rename `convert_request.rs` → `request_translator.rs`
and `stream.rs` → `response_translator.rs` for clarity. Remove `tool_docs.rs` (inlined).
Remove `fallback_to_warp` setting and server fallback path — Bedrock is now the sole backend.
- **Unknown tool handling**: The response translator now detects hallucinated/unknown tool
calls from the model and synthesizes error tool_results so the conversation doesn't
deadlock waiting for a result that will never come.
- **Usage display overhaul**: Replace credit-based usage display with detailed token metrics
showing context window %, cache hit rate (read/write/miss), and estimated cost in dollars.
Add `total_input_tokens`, `total_cache_read_tokens`, `total_cache_write_tokens`, and
`cache_miss_tokens` accessors to `AIConversation`.
- **Predefined rules system**: Add `predefined_rules.rs` with 11 system-defined behavioral
rules that are auto-seeded on first launch. Add "Add Predefined Rules" button to the
Rules UI for re-adding them later. Track seeding state via `has_seeded_predefined_rules`
setting.
- **Session restore improvements**: Rename database file from `warp.sqlite` to
`galaxy.sqlite` with automatic migration from both same-directory and state_dir legacy
paths. Improve CWD persistence by falling back to `session_startup_path` for agent-mode
and fresh tabs. Add extensive session-save/restore logging.
- **Shell bootstrap rebrand**: Rename `WARP_INITIAL_WORKING_DIR` environment variable to
`GALAXY_INITIAL_WORKING_DIR` across bash, zsh, and fish bootstrap scripts.
- **Model defaults**: Change default Bedrock model from Opus 4.7 to Opus 4.6.
Add `context_window_for_model()` helper with model-aware context sizes.
Remove `is_bedrock_model()` (no longer needed without server fallback).
- **User query persistence**: The response translator now emits a `UserQuery` proto message
at stream start so the user's prompt persists across sessions for conversation titles.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Bedrock was calling suggest_next_prompt tool which created a SuggestPrompt
action that waited forever on a oneshot channel for UI interaction that
never fires in the Bedrock path, keeping the conversation permanently
InProgress. Fixed by filtering the tool from the Bedrock tool list and
skipping it at the stream level when the LLM calls it from context history.
Also includes: Bedrock cache token tracking, cost estimation, LSP
improvements, conversation usage view updates, and external config support.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>