Models
Access 38+ leading Chinese AI models through a single API. All models are accessible via the same OpenAI-compatible endpoint.
GET /v1/models to list all available models programmatically. For general tasks, we recommend starting with DeepSeek V4 Pro or Qwen 3.5 397B.
DeepSeek Models
DeepSeek's most capable model. Exceptional at complex reasoning, math, and coding tasks.
Optimized for speed. Great for high-throughput applications with good quality.
Previous generation flagship with proven reliability and broad capability.
Cost-effective general-purpose model. Ideal for most standard applications.
Dedicated reasoning model with chain-of-thought. Excellent for math and logic.
Qwen (Alibaba) Models
Alibaba's next-generation flagship with breakthrough performance.
Massive MoE architecture: 397B total params, 17B active. Outstanding performance.
Battle-tested flagship model. Reliable performance across all tasks.
Best price-performance ratio. Recommended for most production workloads.
Ultra-fast responses for latency-sensitive applications.
Specialized for code generation, review, and debugging across 50+ languages.
Designed for very long documents. Process entire codebases or books.
Image understanding and visual question answering at the highest quality.
Cost-effective image understanding. Great for OCR and image analysis.
Speech and audio understanding. Transcribe and analyze audio inputs.
Other Top Models
Moonshot AI's specialized coding model. Excellent at code generation and review.
Zhipu AI's latest flagship. Strong Chinese language capabilities.
General-purpose model with excellent multilingual performance and creative writing.
StepFun's fast inference model. Good balance of speed and quality.
Meituan's long-context model for processing large documents and codebases.
Professional-grade model with strong analytical and summarization capabilities.
Qwen 3.5 Variant Models
Multiple sizes available for different cost/performance needs:
122B params, 10B active. Good mid-range option.
35B params, 3B active. Ultra-efficient for simple tasks.
27B dense model. Consistent performance for complex tasks.
9B dense model. Fast and economical for lightweight tasks.
4B model. Ultra-fast, lowest cost for simple classification and extraction.
Latest Qwen 3.6 generation. Improved instruction following.
Embedding & Reranking Models
High-quality embedding model for semantic search and RAG. 8B parameters.
Lightweight embedding model. Good quality with lower latency.
Cross-encoder reranking model for improving search result relevance.