OpenAI previews Ultrafast for GPT-5.6 Sol at up to 14x the speed
The new API service tier pairs GPT-5.6 Sol with Cerebras infrastructure to deliver up to 750 output tokens per second for latency-sensitive workloads.
Latest Pubnexa reporting on AI.
The new API service tier pairs GPT-5.6 Sol with Cerebras infrastructure to deliver up to 750 output tokens per second for latency-sensitive workloads.
The new Flash model targets software engineering and multi-step agent workloads while launching at a lower introductory API price.
The 30-billion-parameter open model is built for specialized agent tasks, while NeMo Switchyard routes each step to a suitable model.