Model APIs
Every route, and what it costs.
The whole P/9 catalogue, across both serving pools, priced per million tokens in rupees. One endpoint, one key, one bill.
Global pass-through
Relayed by the gateway to partner capacity outside the PK NPU pool. Same endpoint, same keys, billed in rupees like every other route.
- gpt-oss 120BLive
OpenAI · Global · Chat completions
- Input, Rs / 1M
- Rs 30.08
- Output, Rs / 1M
- Rs 180.47
Route id
paranine/gpt-oss-120b(Global) - Kimi K3Live
Moonshot AI · Global · Chat completions
- Input, Rs / 1M
- Rs 902.34
- Output, Rs / 1M
- Rs 4,601.93
Route id
paranine/Kimi-K3(Global) - GLM-5.3Coming soon
Z.ai · Global · Chat completions
Z.ai's newest model for agentic engineering, with stronger coding and longer-horizon agent work than the 5.2 line.
- Input, Rs / 1M
- Rs 421.09
- Output, Rs / 1M
- Rs 1,323.43
Route id
paranine/GLM-5.3(Global)Proposed, not registered yet
- GLM-5.3-FlashComing soon
Z.ai · Global · Chat completions
Z.ai's first natively multimodal GLM-5 release, tuned for coding and agentic work.
- Input, Rs / 1M
- Rs 45.12
- Output, Rs / 1M
- Rs 150.39
Route id
paranine/GLM-5.3-Flash(Global)Proposed, not registered yet
- GLM-5.2 FastComing soon
Z.ai · Global · Chat completions
GLM-5.2 tuned for speed, for workloads that have to answer in real time.
- Input, Rs / 1M
- Rs 631.64
- Output, Rs / 1M
- Rs 1,985.15
Route id
paranine/GLM-5.2-Fast(Global)Proposed, not registered yet
- Kimi K2.7 CodeComing soon
Moonshot AI · Global · Chat completions
- Input, Rs / 1M
- Rs 285.74
- Output, Rs / 1M
- Rs 1,203.12
Route id
paranine/Kimi-K2.7-Code(Global)Proposed, not registered yet
- DeepSeek-V4-Flash-0731Coming soon
DeepSeek AI · Global · Chat completions
An open-weight mixture of experts, 284B total and 13B active, with a one-million-token context and selectable reasoning effort.
- Input, Rs / 1M
- Rs 39.10
- Output, Rs / 1M
- Rs 78.20
Route id
paranine/DeepSeek-V4-Flash-0731(Global)Proposed, not registered yet
- DeepSeek V4 Pro 0813Coming soon
DeepSeek AI · Global · Chat completions
The 0813 release of DeepSeek's 1.6T-parameter open frontier model.
- Input, Rs / 1M
- Rs 397.03
- Output, Rs / 1M
- Rs 1,191.09
Route id
paranine/DeepSeek-V4-Pro-0813(Global)Proposed, not registered yet
- DeepSeek V4 ProComing soon
DeepSeek AI · Global · Chat completions
- Input, Rs / 1M
- Rs 523.36
- Output, Rs / 1M
- Rs 1,046.72
Route id
paranine/DeepSeek-V4-Pro(Global)Proposed, not registered yet
- Nvidia Nemotron 3 UltraComing soon
Nvidia · Global · Chat completions
- Input, Rs / 1M
- Rs 180.47
- Output, Rs / 1M
- Rs 721.87
Route id
paranine/Nemotron-3-Ultra(Global)Proposed, not registered yet
- InklingComing soon
Thinking Machines Lab · Global · Chat completions
A multimodal mixture of experts, 975B total and 41B active, reasoning over text, images and audio at a 256k context.
- Input, Rs / 1M
- Rs 300.78
- Output, Rs / 1M
- Rs 1,218.16
Route id
paranine/Inkling(Global)Proposed, not registered yet
- Inkling-SmallComing soon
Thinking Machines Lab · Global · Chat completions
The smaller, open-weight Inkling: faster and cheaper to run, with the same native text, image and audio input.
- Input, Rs / 1M
- Rs 150.39
- Output, Rs / 1M
- Rs 360.94
Route id
paranine/Inkling-Small(Global)Proposed, not registered yet
PK NPU-native endpoints (Ascend 910B)
Served from P/9's own Ascend 910B NPU capacity in Pakistan, the platform's native hardware. Open on the same endpoint as every other route, with the same keys and the same rupee billing.
- GLM-5.2Live
Z.ai · PK NPU · Chat completions
The previous generation of Z.ai's agentic engineering line: coding, tool use and sustained execution on long tasks.
- Input, Rs / 1M
- Rs 421.09
- Output, Rs / 1M
- Rs 1,365.54
Route id
paranine/GLM-5.2 - gpt-oss 120BLive
OpenAI · PK NPU · Chat completions
- Input, Rs / 1M
- Rs 30.08
- Output, Rs / 1M
- Rs 180.47
Route id
paranine/gpt-oss-120b(high) - Qwen3.8 27BLive
Qwen team, Alibaba · PK NPU · Chat completions
- Input, Rs / 1M
- Rs 80.00
- Output, Rs / 1M
- Rs 650.00
Route id
paranine/qwen-3.8-27b - Qwen3 Embedding 8BLive
Qwen team, Alibaba · PK NPU · Embeddings
- Input, Rs / 1M
- Rs 2.78
- Output, Rs / 1M
- No output tokens
Route id
paranine/qwen3-embedding
| Model | Route id (the model string) | Input, Rs / 1M | Output, Rs / 1M |
|---|---|---|---|
Global pass-throughRelayed by the gateway to partner capacity outside the PK NPU pool. Same endpoint, same keys, billed in rupees like every other route. | |||
gpt-oss 120BLive OpenAI · Global · Chat completions | paranine/gpt-oss-120b(Global) | Rs 30.08 | Rs 180.47 |
Kimi K3Live Moonshot AI · Global · Chat completions | paranine/Kimi-K3(Global) | Rs 902.34 | Rs 4,601.93 |
GLM-5.3Coming soon Z.ai · Global · Chat completions Z.ai's newest model for agentic engineering, with stronger coding and longer-horizon agent work than the 5.2 line. | paranine/GLM-5.3(Global)Proposed, not registered yet | Rs 421.09 | Rs 1,323.43 |
GLM-5.3-FlashComing soon Z.ai · Global · Chat completions Z.ai's first natively multimodal GLM-5 release, tuned for coding and agentic work. | paranine/GLM-5.3-Flash(Global)Proposed, not registered yet | Rs 45.12 | Rs 150.39 |
GLM-5.2 FastComing soon Z.ai · Global · Chat completions GLM-5.2 tuned for speed, for workloads that have to answer in real time. | paranine/GLM-5.2-Fast(Global)Proposed, not registered yet | Rs 631.64 | Rs 1,985.15 |
Kimi K2.7 CodeComing soon Moonshot AI · Global · Chat completions | paranine/Kimi-K2.7-Code(Global)Proposed, not registered yet | Rs 285.74 | Rs 1,203.12 |
DeepSeek-V4-Flash-0731Coming soon DeepSeek AI · Global · Chat completions An open-weight mixture of experts, 284B total and 13B active, with a one-million-token context and selectable reasoning effort. | paranine/DeepSeek-V4-Flash-0731(Global)Proposed, not registered yet | Rs 39.10 | Rs 78.20 |
DeepSeek V4 Pro 0813Coming soon DeepSeek AI · Global · Chat completions The 0813 release of DeepSeek's 1.6T-parameter open frontier model. | paranine/DeepSeek-V4-Pro-0813(Global)Proposed, not registered yet | Rs 397.03 | Rs 1,191.09 |
DeepSeek V4 ProComing soon DeepSeek AI · Global · Chat completions | paranine/DeepSeek-V4-Pro(Global)Proposed, not registered yet | Rs 523.36 | Rs 1,046.72 |
Nvidia Nemotron 3 UltraComing soon Nvidia · Global · Chat completions | paranine/Nemotron-3-Ultra(Global)Proposed, not registered yet | Rs 180.47 | Rs 721.87 |
InklingComing soon Thinking Machines Lab · Global · Chat completions A multimodal mixture of experts, 975B total and 41B active, reasoning over text, images and audio at a 256k context. | paranine/Inkling(Global)Proposed, not registered yet | Rs 300.78 | Rs 1,218.16 |
Inkling-SmallComing soon Thinking Machines Lab · Global · Chat completions The smaller, open-weight Inkling: faster and cheaper to run, with the same native text, image and audio input. | paranine/Inkling-Small(Global)Proposed, not registered yet | Rs 150.39 | Rs 360.94 |
PK NPU-native endpoints (Ascend 910B)Served from P/9's own Ascend 910B NPU capacity in Pakistan, the platform's native hardware. Open on the same endpoint as every other route, with the same keys and the same rupee billing. | |||
GLM-5.2Live Z.ai · PK NPU · Chat completions The previous generation of Z.ai's agentic engineering line: coding, tool use and sustained execution on long tasks. | paranine/GLM-5.2 | Rs 421.09 | Rs 1,365.54 |
gpt-oss 120BLive OpenAI · PK NPU · Chat completions | paranine/gpt-oss-120b(high) | Rs 30.08 | Rs 180.47 |
Qwen3.8 27BLive Qwen team, Alibaba · PK NPU · Chat completions | paranine/qwen-3.8-27b | Rs 80.00 | Rs 650.00 |
Qwen3 Embedding 8BLive Qwen team, Alibaba · PK NPU · Embeddings | paranine/qwen3-embedding | Rs 2.78 | No output tokens |
How to read this table.
- Coming soon means the tag, not the price
- A route listed ahead of opening answers 404 until the tag comes off. Its rate is published now so an integration can be costed before it ships, and it is the rate the meter will use on day one.
- A proposed id is the shape, not the string
- Some routes carry an id P/9 intends to publish but has not registered on the gateway yet. Those say so under the id. Write against them if you like, but expect to change the string once the route opens.
- Prices are per million tokens, input and output metered apart
- A request is billed on the tokens it actually consumes: the prompt at the input rate, the completion at the output rate. Embedding routes return vectors rather than tokens, so they have no output rate at all.
- There is no listing endpoint yet
- GET /v1/models is not served, so a stock OpenAI SDK's models.list() returns 404. This page, the dashboard's Model APIs panel and a key's allowed_models are the sources of truth until it ships.