Commit Graph

598 Commits

Author SHA1 Message Date
Mario Zechner ef6af5ebbd feat(ai,coding-agent): add faux provider and ModelRegistry factories 2026-03-29 21:08:50 +02:00
Mario Zechner 567249e8b4 Add [Unreleased] section for next cycle 2026-03-29 13:14:04 +02:00
Mario Zechner bc8eb74b82 fix(ai): detect Ollama overflow errors closes #2626 2026-03-27 02:50:22 +01:00
Gordon Hui 17625cc8a2 feat(ai): add google-vertex gemini-3.1-pro-preview-customtools (#2610) 2026-03-27 02:45:58 +01:00
Mario Zechner e24a61ef01 Add [Unreleased] section for next cycle 2026-03-27 02:31:10 +01:00
Mario Zechner 10a02d461b chore(ai): update generated models 2026-03-27 00:47:40 +01:00
Mario Zechner 76f6f8cb8b fix(coding-agent,ai): restore main syntax and apply biome formatting 2026-03-26 00:12:26 +01:00
Mario Zechner 6dc43d6dd1 fix(ai): prune deprecated direct minimax models 2026-03-25 22:44:32 +01:00
sparkleMing 6d744f02ef fix: subtract cached tokens from input in Google and Vertex cost calculation (#2588)
Google's promptTokenCount includes cachedContentTokenCount, so using it
directly as the input token count causes double-counting when calculateCost
multiplies input by the input rate AND cacheRead by the cacheRead rate.

The google-gemini-cli provider already handles this correctly (subtracting
cachedContentTokenCount from promptTokenCount), but google.ts and
google-vertex.ts were using the raw promptTokenCount.

This fix aligns both providers with the google-gemini-cli behavior.
2026-03-25 22:02:28 +01:00
Mario Zechner bab58f821d fix(ai): omit copilot responses reasoning default closes #2567 2026-03-24 20:39:11 +01:00
Mario Zechner 05c17cfbfe Add [Unreleased] section for next cycle 2026-03-23 02:50:43 +01:00
Mario Zechner d1613e3f53 fix(ai): handle explicit thinking off across providers closes #2490 2026-03-22 20:27:44 +01:00
Mario Zechner 6129971c04 fix(ai): explicitly disable Anthropic thinking when off closes #2022 2026-03-22 19:38:54 +01:00
wjonaskr 3bcbae490c feat(ai): add requestMetadata support to BedrockOptions for cost allocation tagging (#2511)
Add an optional requestMetadata field to BedrockOptions that forwards
key-value pairs to the Bedrock Converse API ConverseStreamCommand. Tags
appear in AWS Cost Explorer split cost allocation data, enabling callers
to attribute inference costs to specific applications or contexts.

Changes:
- Add requestMetadata?: Record<string, string> to BedrockOptions with
  JSDoc documenting AWS constraints (max 50 pairs, key 64 chars, value
  256 chars, no aws: prefix)
- Pass requestMetadata to commandInput via conditional spread to avoid
  sending undefined when omitted
- Export BedrockOptions from package root (consistent with other
  provider option types)
- Add E2E tests: metadata forwarded to SDK payload, and omitted when
  not provided

closes #2510
2026-03-22 19:05:55 +01:00
Mario Zechner b21b42d032 fix(ai): hash foreign responses tool item ids 2026-03-22 05:40:26 +01:00
Mario Zechner 60358dc493 Merge remote-tracking branch 'origin/main' 2026-03-20 20:16:07 +01:00
Mario Zechner a84fd4e3cf Add [Unreleased] section for next cycle 2026-03-20 20:15:24 +01:00
Cheng-Zi-Qing 7e2689ac18 packages/ai: ignore null chunks in openai-completions streams (#2466) 2026-03-20 20:01:28 +01:00
简简简简 8705fbee54 fix(models): align minimax and zai defaults (#2445)
Update the coding-agent default model picks for ZAI, Cerebras,
and MiniMax so new sessions prefer the current model lineup.

Add the missing MiniMax-M2.1-highspeed direct provider entries
and normalize MiniMax Anthropic-compatible context limits so the
catalog matches the provider's supported model set.
2026-03-20 10:00:57 +01:00
Mario Zechner 527c2af5e7 chore(ai): normalize MiMo V2 Pro model name 2026-03-20 01:55:30 +01:00
Mario Zechner 0d7c81ec9e fix(ai): skip AJV validation in restricted runtimes closes #2395 2026-03-19 22:51:12 +01:00
Mario Zechner ac5e0ef76e Updated models 2026-03-19 22:13:27 +01:00
Mario Zechner ff1ea12324 fix(ai): ignore placeholder vertex api keys closes #2335 2026-03-18 22:51:14 +01:00
PriNova f704ee7255 fix(ai): use OpenRouter reasoning payload (#2298)
* fix(ai): use OpenRouter reasoning payload

* fix(coding-agent): stop updating packages on startup closes #1963

* fix(ai): add openrouter thinkingFormat compat

---------

Co-authored-by: Mario Zechner <badlogicgames@gmail.com>
2026-03-18 11:23:27 +01:00
Jheng-Hong Yang 68da22f18c feat(ai): add openai-codex gpt-5.4-mini (#2334) 2026-03-18 11:22:11 +01:00
xu0o0 31d59f8513 fix(ai): support prompt caching for Bedrock application inference profiles (#2346)
Add AWS_BEDROCK_FORCE_CACHE=1 environment variable support. When the
model ID doesn't contain a recognizable Claude model name, users can
set this variable to force cache point injection.
2026-03-18 11:21:10 +01:00
Mario Zechner 6a95f6882c Add [Unreleased] section for next cycle 2026-03-18 03:42:13 +01:00
Mario Zechner b2548ce489 fix(ai): normalize replayed responses tool call ids closes #2328 2026-03-18 03:37:40 +01:00
Mario Zechner 453b22397d fix(ai): keep image tool results inline for gemini 3+ and antigravity closes #2052 2026-03-18 02:17:44 +01:00
Mario Zechner d70dfbeb3e fix(ai): correct Bedrock Claude 4.6 context window to 200k
Bedrock Claude Opus 4.6 and Sonnet 4.6 models have 200k context
window, not 1M. Removed incorrect overrides that were forcing these
models to 1M. The native Anthropic API models correctly remain at 1M.

closes #2305
2026-03-18 00:08:50 +01:00
Mario Zechner 8dc2bb969c fix(coding-agent,ai): restore Bun binary lazy provider loading closes #2314 2026-03-17 23:30:18 +01:00
Mario Zechner 94ba13ccef fix(ai): align oauth callback flows closes #2316 2026-03-17 23:13:44 +01:00
Mario Zechner 3563cc4df6 fix(ai): resolve codex oauth callback immediately closes #2316 2026-03-17 22:47:51 +01:00
Mario Zechner d914d1c199 fix(coding-agent): handle z.ai network_error closes #2313 2026-03-17 22:34:49 +01:00
Mario Zechner b221d10bd8 Add [Unreleased] section for next cycle 2026-03-17 18:13:12 +01:00
Mario Zechner a7559f01e9 feat(ai): lazy-load provider modules for faster startup fixes #2297 2026-03-17 18:03:24 +01:00
Mario Zechner dd53eb56ee fix(ai): expose provider response ids on assistant messages fixes #2245 2026-03-17 17:11:38 +01:00
Mario Zechner 635bda80d4 Merge commit 0acb65fe from PR #2058 2026-03-17 13:31:00 +01:00
Mario Zechner 40a1300089 Add [Unreleased] section for next cycle 2026-03-16 20:29:18 +01:00
Mario Zechner aac0e0c1bf Add [Unreleased] section for next cycle 2026-03-15 20:34:10 +01:00
Mario Zechner e645995a4a fix(ai): fix anthropic oauth manual callback and refresh flow closes #2169 2026-03-15 15:40:50 +01:00
Mario Zechner ad48b52de4 fix(ai): align codex websocket headers and terminate SSE closes #1961 2026-03-14 12:23:17 +01:00
Mario Zechner 1a4d153d7a fix(ai): limit Bedrock prompt caching to Claude models\n\ncloses #2053 2026-03-14 05:20:41 +01:00
Mario Zechner 03ad7ecee7 fix(ai): add qwen-chat-template compat mode closes #2020 2026-03-14 05:20:27 +01:00
Mario Zechner c961cda2cd fix(ai): harden Bedrock unsigned thinking replay closes #2063 2026-03-14 05:06:14 +01:00
Mario Zechner 62064b2b64 fix(ai): support xhigh for opus 4.6 by model id closes #2040 2026-03-14 04:55:03 +01:00
Mario Zechner a79ca41190 fix(ai): handle unknown finish_reason in openai-completions gracefully
Map unknown finish_reason values (e.g. "end" from Ollama/LM Studio) to
"stop" instead of throwing, since assistant content is already produced.

fixes #2142
2026-03-14 04:01:22 +01:00
Mario Zechner c7309aedac fix(ai): replace curl with fetch in Anthropic OAuth token exchange 2026-03-14 03:21:08 +01:00
Mario Zechner 9b794558c6 chore(ai): refresh generated models 2026-03-14 03:11:15 +01:00
Mario Zechner 933348593c fix(ai): send tool result images in function_call_output closes #2104 2026-03-14 00:49:40 +01:00