AIHubMix Blog

Announcements, tutorials, model analysis, and product updates from AIHubMix.

Migrating to GPT-6.1 Sol: 9 Things That Can Go Wrong

Migrating to GPT-6.1 Sol: 9 Things That Can Go Wrong

Moving from GPT-6 Sol to 6.1 Sol looks like a one-line change, and the price is the same. But there are a few breaking changes and some shifts in behavior, so changing only the model name can get you errors, a surprise bill, or an agent that behaves differently. OpenAI's GPT-6 migration guide covers most of the official changes. This post adds the things that tend to bite in practice. They're ordered from "fails loudly" to "fails quietly." 1. reasoning_effort: "none" returns a 400 What you'l

8 min readTutorial
What GPT-6.1 Sol Really Costs: Beyond the $2 / $10 Price Tag

What GPT-6.1 Sol Really Costs: Beyond the $2 / $10 Price Tag

GPT-6.1 Sol has the same list price as GPT-6 Sol: $2 per million input tokens and $10 per million output tokens. Your actual bill depends on four other things: how often you hit the cache, whether you cross 272K input tokens, which service tier you use, and how many reasoning tokens the model burns. This post goes through each one. The full price list Standard rates (per 1M tokens) GPT-6.1 SolGPT-6 SolGPT-6 AstraGPT-6 LunaInput$2.00$2.00$10.00$0.10Cached input$0.10$0.20$1.00—Cache writes$

6 min readOpinion
Choosing a Reasoning Effort for GPT-6.1 Sol: low to max

Choosing a Reasoning Effort for GPT-6.1 Sol: low to max

GPT-6.1 Sol has five reasoning effort levels: low, medium (the default), high, xhigh, and max. The big change from GPT-6 Sol is that none and minimal are gone, so low is now the floor. This one setting drives both latency and cost. You never see reasoning tokens, but you pay for them at the output rate ($10 per million for 6.1 Sol on AIHubMix), and they take up room in the context window. Pick the wrong level and you can easily pay several times more than you need to. What each level is for

6 min readTutorial
GPT-6.1 Sol vs GPT-6 Sol: A One-Week Upgrade That Nearly Catches Astra

GPT-6.1 Sol vs GPT-6 Sol: A One-Week Upgrade That Nearly Catches Astra

OpenAI shipped GPT-6.1 Sol at DevDay on September 29, 2026. The timing is striking: GPT-6 Sol had been out for exactly one week (it launched September 22 alongside Luna), and the flagship GPT-6 Astra was less than a month old. The official model page sums up the pitch in one line: "Near-Astra performance for complex work at a lower cost." Concretely, that means agentic coding, computer use, and professional tasks at roughly Astra quality, for one fifth of Astra's per-token price. This post ans

7 min readNews
AIHubMix Adds Claude Sonnet 5.5: An Upgrade for Everyday Work at Unchanged Token Prices

AIHubMix Adds Claude Sonnet 5.5: An Upgrade for Everyday Work at Unchanged Token Prices

A 1-million-token context window, input at $2 and output at $10 per million tokens, with browser access and a shared gateway to multiple model providers. AIHubMix now offers Claude Sonnet 5.5, Anthropic's latest Sonnet model. Users can access it through AIHubMix Playground or compatible applications for writing, organizing information, preparing presentation content, and other everyday work. Compared with Sonnet 5, Anthropic reports faster output generation, clearer writing, and more efficient

3 min read
Xiaomi MiMo v2.6 Pro Quick Start: Three Steps from Browser to Code

Xiaomi MiMo v2.6 Pro Quick Start: Three Steps from Browser to Code

The quick answer: open https://playground.aihubmix.com/ in your browser, type a question, press send. No installing, no coding, and free trial calls without entering a card. Key points * The model is mimo-v2.6-pro, made by Xiaomi, released 22 September 2026 * You can use it in a browser for free before deciding anything * It reads text, images, video and audio * It costs $0.48 per million input tokens, $0.96 per million output — a fraction of a cent for a typical question * The older mim

4 min read
How to Use Xiaomi MiMo v2.6 Pro on AIHubMix: Setup Guide

How to Use Xiaomi MiMo v2.6 Pro on AIHubMix: Setup Guide

In short: point the official openai Python SDK at https://aihubmix.com/v1, use model ID mimo-v2.6-pro, and authenticate with your AIHubMix key. The API is OpenAI-compatible (callable in the same format as OpenAI's), so no special client is required. Setup takes about five minutes. Key points * OpenAI-compatible API — change base_url, nothing else * Paid mimo-v2.6-pro · free xiaomi-mimo-v2.6-pro-free [1] [2] * $0.48/M input · $0.96/M output [1] * Free tier 5 requests/min · 100/day · 1M tok

6 min read
How to Call Jev on AiHubMix: A Structured Classification Tutorial

How to Call Jev on AiHubMix: A Structured Classification Tutorial

Short answer: jev-1.13 does not generate text. You POST a piece of text plus a set of named questions to https://aihubmix.com/v1/systemone, and you get back typed answers keyed by your question names — a category, a score, or a probability. Nothing to parse. Below is a working request, the full response format, and the one design mistake that will quietly cost you accuracy. What Jev is for Use it when you need a judgment, not a paragraph: routing a support ticket, scoring severity, flagging u

5 min read
DeepSeek V4.1 Flash API Pricing Compared: AIHubMix 30% Off vs OpenRouter

DeepSeek V4.1 Flash API Pricing Compared: AIHubMix 30% Off vs OpenRouter

DeepSeek V4.1 Flash is now available on AIHubMix and OpenRouter. It supports text and image inputs, a 1M-token context window, tool calling, structured outputs, and agent-oriented workflows. AIHubMix currently offers a limited-time 30% discount on selected provider routes, valid through September 27, 2026. OpenRouter lists the price at $0.13 per million input tokens and $0.52 per million output tokens, with cache reads at $0.0026 per million tokens. DeepSeek V4.1 Flash API pricing at a glance

3 min read
Why did Gemini turn Chinese for no reason?

Why did Gemini turn Chinese for no reason?

Gemini may occasionally produce Chinese in its reasoning or final response even when you expect English. This usually points to temporary multilingual language drift. It does not, by itself, mean that Gemini has switched to a separate Chinese model. The short version: inspect the full conversation for Chinese language signals, restate the desired output language, and do not use the stream setting as a language fix. What Gemini language drift means Language drift occurs when a model begins in

3 min read
Jev Explained: How to Add Fast, Typed Decisions to an AI Agent

Jev Explained: How to Add Fast, Typed Decisions to an AI Agent

Jev is best understood as a decision layer for software. It reads text or structured state and returns predefined classifications, scores, and yes-or-no probabilities. It does not write an answer for the user. That narrower interface makes it relevant to high-volume routing, triage, verification, and guardrail steps inside AI agents. The practical pattern is simple: let Jev make frequent, reversible judgments; let business code enforce policy; escalate uncertain or consequential cases to a capa

6 min read
How to Create an AI Product Ad with Claude Code and AIHubMix

How to Create an AI Product Ad with Claude Code and AIHubMix

With one clear product image, Claude Code, and the AIHubMix API, you can generate character references in Seedream, register them as virtual-person assets, and use Seedance to create an AI product ad. This tutorial is based on a real aloe-mist campaign and includes the project structure, configuration, commands, and complete video prompt. The target output is a 20-second, 720p, 9:16 vertical ad with an AI-generated person and setting. Claude Code and AIHubMix power the full workflow. Claude Co

9 min read
Amp Reopens BYOK After a Year: Connecting and Configuring AIHubMix

Amp Reopens BYOK After a Year: Connecting and Configuring AIHubMix

Amp has reopened BYOK (Bring Your Own Key), and AIHubMix now supports it. This guide uses Anthropic's Claude models as the example and walks through pointing Amp at the AIHubMix API via Custom URL and Model Routing — covering the API key, model name mapping, and project setup. If you're searching for "Amp BYOK setup", "connect Amp to AIHubMix", "AIHubMix API key", "Amp Custom URL", or "Amp Model Routing", the steps below cover it. What is Amp good for? * Multi-model workflows: assign differ

4 min read
How to Use a Real Face in Seedance 2.5 with AIHubMix Real-Human Assets

How to Use a Real Face in Seedance 2.5 with AIHubMix Real-Human Assets

AI video becomes much more useful when the person on screen stays recognizable from one scene to the next. AIHubMix real-human assets provide a consent-based way to use an approved face in Seedance video generation while keeping authorization, asset processing, and generation as separate, verifiable steps. This guide explains what real-human video is, how the AIHubMix workflow works, why it is useful, and how to create a Seedance 2.5 video from a real face. It also explains the difference betwe

5 min readTutorial
GLM-5.3-Flash Pricing Compared: OpenRouter, Z.ai, and AIHubMix

GLM-5.3-Flash Pricing Compared: OpenRouter, Z.ai, and AIHubMix

GLM-5.3-Flash pricing on OpenRouter, Z.ai, and AIHubMix, including OpenRouter's 5.5% platform fee and cache-ratio cost examples.

4 min readNews
AI Agent Architecture: Model Routing and Tool Discovery

AI Agent Architecture: Model Routing and Tool Discovery

Learn why production AI agents need two integrations: AIHubMix for real-time model routing and Monid for runtime tool discovery and API access.

8 min readOpinion
Building the Agent Economy: AIHubMix and FluxA Announce Strategic Collaboration

Building the Agent Economy: AIHubMix and FluxA Announce Strategic Collaboration

At AIHubMix, we believe the next phase of artificial intelligence will be defined not only by what models can understand, but also by what AI agents can responsibly accomplish. That is why we are announcing a strategic collaboration with FluxA, an agent-native payment infrastructure provider focused on wallets, identity, authorization, service monetization, and payments for autonomous systems. Together, AIHubMix and FluxA are working to connect two essential layers of the Agent Economy: access

3 min readNews
Tell Your Agent One Sentence, Get 800+ Models"

Tell Your Agent One Sentence, Get 800+ Models"

Agent-readable access on the AIHubMix AI gateway: agents.md, llms.txt for 850+ models, an Agent Skill, MCP, and deep links for Claude Code, Codex, Cursor.

5 min readAnnouncement
GPT-5.6 Sol API Pricing: 50% Off on AIHubMix and OpenRouter

GPT-5.6 Sol API Pricing: 50% Off on AIHubMix and OpenRouter

GPT-5.6 Sol is currently 50% off on AIHubMix and OpenRouter, but OpenRouter adds a 5.5% Pay-as-you-go platform fee. Compare the real final cost.

5 min readNews
GLM-5.3 Pricing Compared: OpenRouter, Z.ai, and AIHubMix

GLM-5.3 Pricing Compared: OpenRouter, Z.ai, and AIHubMix

GLM-5.3 has quickly become one of the most interesting models for coding and long-horizon agent work. But the price you pay depends heavily on where you access it. The official Z.ai API and OpenRouter currently list the model at $1.40 per million input tokens and $4.40 per million output tokens. AIHubMix offers a separate coding-glm-5.3 preview route at $0.06 per million input tokens and $0.22 per million output tokens. That is a dramatic difference. It is also not a simple apples-to-apples co

3 min readNews
GLM-5.3 Hands-on Guide: Always-on Thinking, Three Effort Levels, and the API Support Matrix

GLM-5.3 Hands-on Guide: Always-on Thinking, Three Effort Levels, and the API Support Matrix

An August 2026 guide to calling GLM-5.3: always-on thinking with three reasoning_effort levels, reasoning summaries, parallel tool calls, structured output, and automatic caching — with tested examples for the AIHubMix Chat, Responses, and Messages APIs.

7 min readTutorial
DeepSeek V4 Pro (0813): Thinking Passback & 3-API Matrix

DeepSeek V4 Pro (0813): Thinking Passback & 3-API Matrix

DeepSeek V4 Pro (0813) hands-on guide: thinking toggle and reasoning_effort levels, mandatory thinking-history passback, tools, caching, and a 3-API matrix.

15 min readTutorial
DeepSeek V4 Flash Was Degraded Today. Here’s Why Multi-Provider Failover Matters

DeepSeek V4 Flash Was Degraded Today. Here’s Why Multi-Provider Failover Matters

On August 4, 2026, DeepSeek’s official status page recorded two API degraded-performance incidents. The first incident lasted 1 hour and 18 minutes, from 02:02 to 03:20 UTC, and affected DeepSeek V4 Flash, V4 Pro, and Expert Mode. The second incident lasted 36 minutes, from 03:43 to 04:20 UTC, and affected the DeepSeek V4 Flash API. Both incidents have since been resolved. OpenCode also reported that DeepSeek Flash was experiencing capacity issues due to unprecedented demand. However, DeepSeek

4 min readOpinion
June 2026 Release Spotlight: ~20 New Models

June 2026 Release Spotlight: ~20 New Models

In June 2026 AIHubMix added ~20 models (glm-5.2, minimax-m3, qwen3.7-plus, kimi-k2.7-code, Kling video) plus LLM Router, Mapping & Fallback, CLI, backup domain.

3 min readChangelog
Global Acceleration: 75% Lower Latency, 99.99% Availability

Global Acceleration: 75% Lower Latency, 99.99% Availability

AIHubMix runs a self-built acceleration network: 75% lower latency, 60% less fluctuation, 99.99% availability, minute-level probes, automatic failover.

2 min readAnnouncement
OpenAI Compatible Interface Upgrade: Deep Support for Claude

OpenAI Compatible Interface Upgrade: Deep Support for Claude

AIHubMix upgraded its OpenAI-compatible API for Claude: interleaved thinking with no extra parameters, prompt caching, and Anthropic beta feature support.

8 min readNews
GPT-5.6 Is Live: Prompt Caching Billing Changes Explained

GPT-5.6 Is Live: Prompt Caching Billing Changes Explained

July 2026: GPT-5.6 gpt-5.6-sol / terra / luna on AIHubMix: 1.05M context, 1.25x cache writes, prompt_cache_key, explicit breakpoints, vs Claude caching.

8 min readOpinion
July 2026 Release Spotlight: ~30 New Models and Media APIs

July 2026 Release Spotlight: ~30 New Models and Media APIs

AIHubMix added ~30 models in July 2026, including claude-opus-5, GPT-5.6, kimi-k3 and qwen3.8-max-preview, plus media generation, 3D generation and MCP Server.

6 min readChangelog
Free AI Models on AIHubMix

Free AI Models on AIHubMix

Free AI Models: The Ultimate Guide to Building with Zero-Cost AI in 2026 on AIHubMix

13 min readAnnouncement
Kimi K3 Hands-On Guide: New Parameters & API Support Matrix

Kimi K3 Hands-On Guide: New Parameters & API Support Matrix

July 2026 Kimi K3 guide: reasoning_effort max, thinking history, dynamic tool loading, structured output, auto caching, partial prefix, and vision inputs.

8 min readTutorial
Claude Opus 4.7 New Parameters Guide

Claude Opus 4.7 New Parameters Guide

This article covers two key changes to reasoning control in Claude Opus 4.7, along with complete usage instructions for both the AIHubmix native API and the Chat unified interface. See also: Anthropic official announcement and model change log.

3 min readTutorial