What are the main differences between GPT-6.1 Sol and GPT-6 Sol?
Key official release highlights include improvements in real codebase tasks, complex PDF analysis, and multi-step business workflows, as well as better factual accuracy and lower official cached input pricing. Similar capacity does not mean calls are fully compatible; tool requests and reasoning tiers need to be verified for the new model.
Can all 1.05M tokens be used for input?
No. The official model documentation separately lists a maximum input of 922,000 tokens and a maximum output of 128,000 tokens; the 1,050,000-token context must accommodate both relevant input and output. Actual requests must also meet this platform's API limits, and images, tool content, and history must be included in the budget.
Which API should be used for tool calling?
Tool workflows use Responses, with function definitions, call results, and execution information configured and returned according to the platform documentation. The latest official model documentation explicitly states that Chat Completions for gpt-6.1-sol does not support tool calling; support cannot be inferred merely because shared requests include a tools field. Standard text-and-image Q&A can use Chat Completions.
How should reasoning effort be selected?
The platform guide recommends starting with low, comparing medium, high, or xhigh using the same set of tasks, then evaluating max for complex tasks that genuinely require it. Higher tiers should be determined by completion quality, while also considering reasoning usage and latency. Do not continue using none or minimal, and set reasoning_effort or reasoning according to the selected API.
How do caching and long context affect billing?
This platform's billing rules distinguish among standard input, cached input reads, cached input writes, and output; when prompt_tokens exceeds 272,000, the long-context tier applies, with input and cache-related rates at 2 times the corresponding base tier and output at 1.5 times. Amounts are subject to the Pricing page and selected plan; repeated text does not necessarily mean the cache has been hit.