Tale 0.5.72 runs GPT-6 Astra tool calls through the Responses API

Tale 0.5.72 routes GPT-6 Astra tool calls through OpenAI's Responses API and restricts tasks using it to Codex, while subscriptions remain task-only.

Tale 0.5.72 routes tool calls for OpenAI's GPT-6 Astra and GPT-6.1 Sol through the Responses API (POST /responses). Their Chat Completions support is text-only for this purpose. In tasks and automations, Tale runs both models exclusively on Codex, the only bundled harness with Responses API wire support. Direct chat requires an OpenAI API key or environment variable, while ChatGPT subscription credentials remain restricted to tasks and automations.

Update on 2026-10-05: OpenAI's changelog lists a 2026-09-29 Ultrafast-mode addition for GPT-6 Astra in the Responses API. That adds a service-tier option, but it does not change Tale 0.5.72's tool-calling route.

Key facts

  • On 2026-10-04, Tale v0.5.72 added GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, and GPT-6.1 Sol to chat; Astra and GPT-6.1 Sol also run in project agents and automations on Codex.
  • GPT-6 Astra and GPT-6.1 Sol support Chat Completions for plain text, but Tale uses the Responses API for their function calls.
  • Tale's catalog gives GPT-6 Astra and GPT-6.1 Sol no off setting; GPT-6 Sol and GPT-6 Luna declare that tools require reasoning set to none.
  • Tale constructs a dedicated Responses API payload in chat_wire.ts using top-level instructions, input items, store: false, and strict: false function definitions.
  • Project agents and automations require the Codex harness for Responses-only models; configuring any other runtime triggers an explicit refusal error.
  • Vendor subscription credentials (keys and brokers) operate only in tasks and automations; direct chat excludes them because vendors restrict subscription tokens to their proprietary runtimes.
  • Tale's external Chat Completions proxy (/api/v1/openai/...) omits GPT-6 Astra and GPT-6.1 Sol because the proxy endpoint does not relay the Responses API.

How Tale 0.5.72 shapes Responses API calls

Tale's v0.5.72 release notes, published on 2026-10-04, say that GPT-6 Astra and GPT-6.1 Sol support Chat Completions for plain text but require the Responses API (/responses) for function calling. Pull request #4181 merged on 2026-10-03 and added Responses handling in chat_wire.ts.

When a model's catalog entry specifies toolCallingApi: responses, Tale sends an HTTP POST request to <baseUrl>/responses instead of <baseUrl>/chat/completions. The request transforms chat history into the schema expected by the Responses API:

  1. System instructions: System messages are extracted and passed as the top-level instructions parameter rather than inside the conversation array.
  2. Conversation items: The transcript is mapped to an array of input items. User turns become message objects, assistant responses become text items followed by function_call items, and tool execution outputs are passed as function_call_output items linked by call_id.
  3. Tool definitions: Function schemas are offered with strict: false. Tale applies this flag because its existing tool definitions were authored for Chat Completions, whereas the Responses API enforces strict schema validation by default.
  4. Output caps and effort: Completion limits use max_output_tokens. Tale maps seven internal effort choices (none, minimal, low, medium, high, extra, and max) to Responses values; extra becomes xhigh.
  5. Data retention: Every request sets store: false to ensure OpenAI does not persist conversation state server-side.

On the streaming side, Tale's SSE decoder in stream_decode.ts parses incoming events into chunks:

switch (type) {
  case 'response.output_text.delta':
  case 'response.refusal.delta':
    return { text: delta };
  case 'response.reasoning_summary_text.delta':
  case 'response.reasoning_text.delta':
    return { text: '', ...(delta ? { reasoning: delta } : {}) };
  case 'response.output_item.added':
  case 'response.output_item.done':
    // Accumulates tool call id and name into state.drafts
    return { text: '' };
  case 'response.function_call_arguments.delta':
    // Drips argument fragments into state.drafts
    return { text: '' };
  case 'response.function_call_arguments.done':
    // Replaces fragments with the complete argument string
    return { text: '' };
  case 'response.completed':
  case 'response.incomplete':
  case 'response.failed':
    // Emits final token usage and finish reason
    return { text: '', usage: totals(state.running), finishReason };
}

The v0.5.72 release notes state that GPT-6 models in chat do not carry reasoning across tool rounds within a turn. Tale sends the transcript as input items on each Responses API request instead of relying on stored conversation state.

Model configuration in the Tale catalog

Tale 0.5.72 cataloged the four GPT-6 models in configs/platform/system/models/openai/models.yml:

Model ID

Tool API

Reasoning knob

Context window

Output limit

Input / Output (per 1M tokens)

gpt-6-astra

Responses

effort (no off setting)

1,050,000

128,000

$10.00 / $50.00

gpt-6.1-sol

Responses

effort (no off setting)

1,050,000

128,000

$2.00 / $10.00

gpt-6-sol

Chat Completions

effort (tools require none)

1,050,000

128,000

$2.00 / $10.00

gpt-6-luna

Chat Completions

effort (tools require none)

1,050,000

128,000

$0.10 / $0.50

Because GPT-6 Sol and GPT-6 Luna require reasoning_effort: none when calling functions over Chat Completions, Tale's chat model picker removes the reasoning effort control whenever tools are active for those two models.

Why tasks run GPT-6 Astra only on Codex

Among these runtimes, only Codex speaks the openai-responses wire protocol. The execution resolver (resolve_execution.ts) enforces this compatibility check via supportsToolCallingWire:

export function supportsToolCallingWire(
  model: Pick<ModelCatalogEntry, 'toolCallingApi'>,
  wire: HarnessGatewayWire,
): boolean {
  return model.toolCallingApi !== 'responses' || wire === 'openai-responses';
}

If an operator assigns GPT-6 Astra or GPT-6.1 Sol to an agent running on another harness (such as OpenCode or Claude Code), agent_serving.ts aborts execution with an explicit error:

model "gpt-6-astra" takes tools only through the Responses API, which the "opencode" harness does not speak — run the agent on "codex", or pick another model

In the admin interface under Settings > AI providers, the Agent runtimes section reports no-compatible-model for any harness that lacks a direct model matching its supported wire, preventing misleading reports of missing credentials.

Why subscription credentials stay out of direct chat

Tale supports two authentication methods for provider subscriptions: static subscription keys and dynamic subscription brokers. These credentials support Claude Code for Anthropic subscriptions and Codex for OpenAI ChatGPT subscriptions.

Vendor subscriptions are restricted entirely to tasks and automations:

  • Vendor licensing terms: Upstream providers forbid utilizing subscription tokens in external applications. Anthropic actively rejects subscription authorization headers from clients other than Claude Code.
  • Isolated sandbox execution: In tasks and automations, the vendor runtime runs within a containerized sandbox that hosts the vendor's official binary and CLI session. Direct chat runs directly on the Tale platform backend, where vendor subscription tokens cannot be utilized.
  • Interface disclosures: Tale marks subscription credentials as Tasks and automations only in the credential configuration dialog and table. In the chat composer, the model dropdown lists the specific providers omitted from chat due to subscription constraints.

To use GPT-6 Astra or GPT-6.1 Sol in direct chat, organizations must provide an OpenAI API key or configure a deployment environment variable prefixed with TALE_PROVIDER_KEY_.

Sources

Last verified: 2026-10-06.

Spotted an outdated or wrong claim? Agents can report it with evidence throughPOST /api/feedback; an editor checks every report. See llms.txt for the agent API.