Models & limits
What you can call, how much, and which parameters work.
Models
| Fact | Qwen3.5 9B |
|---|---|
| Model id | qwen3.5-9b |
| Vendor | Qwen (Alibaba) |
| Parameters | 9B |
| Quantization | Q4_K_M |
| Context window | 16,384 tokens (prompt + reply) |
| Max output | 2,048 tokens |
| Input | Text |
| Capabilities | Chat, Streaming, Tool calling, JSON mode |
| License | Apache-2.0 |
| Price | Free during beta |
| Hosting | Spain (EU) |
GET /v1/models returns the same list as JSON.
Limits
Beta limits apply per API key.
| Limit | Value | When you hit it |
|---|---|---|
| Requests per minute | 60 | 429 rate_limit_requests |
| Tokens per day | 500,000 | 429 rate_limit_tokens. Resets at 00:00 UTC. |
| Concurrent requests | 2 | 429 concurrency_limit |
Every accepted request reports where you stand in rate-limit headers.
Supported
- Chat:
POST /v1/chat/completionswithsystem,user,assistantandtoolmessages.developeris treated assystem. - Streaming:
stream: true. Addstream_options.include_usageto get token counts in the last chunk. - Tools:
tools,tool_choice,parallel_tool_calls. - JSON output:
response_formatwith{"type": "json_object"}. - Sampling:
temperature,top_p,stop,seed,presence_penalty,frequency_penalty. - Length:
max_tokensormax_completion_tokens, up to 2,048. - Also passed through:
logprobs,top_logprobs,top_k,min_p,repeat_penalty.
Not supported
These return 400 with code unsupported_parameter or unsupported_content:
ngreater than 1logit_biasaudio, andmodalitiesother than text- Images or other non-text content in messages
Thinking mode is off. The model answers directly.
Any other parameter is ignored.