Model APIs

Every route, and what it costs.

The whole P/9 catalogue, across both serving pools, priced per million tokens in rupees. One endpoint, one key, one bill.

Global pass-through

Relayed by the gateway to partner capacity outside the PK NPU pool. Same endpoint, same keys, billed in rupees like every other route.

  • gpt-oss 120BLive

    OpenAI · Global · Chat completions

    Input, Rs / 1M
    Rs 30.08
    Output, Rs / 1M
    Rs 180.47

    Route id

    paranine/gpt-oss-120b(Global)
  • Kimi K3Live

    Moonshot AI · Global · Chat completions

    Input, Rs / 1M
    Rs 902.34
    Output, Rs / 1M
    Rs 4,601.93

    Route id

    paranine/Kimi-K3(Global)
  • GLM-5.3Coming soon

    Z.ai · Global · Chat completions

    Z.ai's newest model for agentic engineering, with stronger coding and longer-horizon agent work than the 5.2 line.

    Input, Rs / 1M
    Rs 421.09
    Output, Rs / 1M
    Rs 1,323.43

    Route id

    paranine/GLM-5.3(Global)

    Proposed, not registered yet

  • GLM-5.3-FlashComing soon

    Z.ai · Global · Chat completions

    Z.ai's first natively multimodal GLM-5 release, tuned for coding and agentic work.

    Input, Rs / 1M
    Rs 45.12
    Output, Rs / 1M
    Rs 150.39

    Route id

    paranine/GLM-5.3-Flash(Global)

    Proposed, not registered yet

  • GLM-5.2 FastComing soon

    Z.ai · Global · Chat completions

    GLM-5.2 tuned for speed, for workloads that have to answer in real time.

    Input, Rs / 1M
    Rs 631.64
    Output, Rs / 1M
    Rs 1,985.15

    Route id

    paranine/GLM-5.2-Fast(Global)

    Proposed, not registered yet

  • Kimi K2.7 CodeComing soon

    Moonshot AI · Global · Chat completions

    Input, Rs / 1M
    Rs 285.74
    Output, Rs / 1M
    Rs 1,203.12

    Route id

    paranine/Kimi-K2.7-Code(Global)

    Proposed, not registered yet

  • DeepSeek-V4-Flash-0731Coming soon

    DeepSeek AI · Global · Chat completions

    An open-weight mixture of experts, 284B total and 13B active, with a one-million-token context and selectable reasoning effort.

    Input, Rs / 1M
    Rs 39.10
    Output, Rs / 1M
    Rs 78.20

    Route id

    paranine/DeepSeek-V4-Flash-0731(Global)

    Proposed, not registered yet

  • DeepSeek V4 Pro 0813Coming soon

    DeepSeek AI · Global · Chat completions

    The 0813 release of DeepSeek's 1.6T-parameter open frontier model.

    Input, Rs / 1M
    Rs 397.03
    Output, Rs / 1M
    Rs 1,191.09

    Route id

    paranine/DeepSeek-V4-Pro-0813(Global)

    Proposed, not registered yet

  • DeepSeek V4 ProComing soon

    DeepSeek AI · Global · Chat completions

    Input, Rs / 1M
    Rs 523.36
    Output, Rs / 1M
    Rs 1,046.72

    Route id

    paranine/DeepSeek-V4-Pro(Global)

    Proposed, not registered yet

  • Nvidia Nemotron 3 UltraComing soon

    Nvidia · Global · Chat completions

    Input, Rs / 1M
    Rs 180.47
    Output, Rs / 1M
    Rs 721.87

    Route id

    paranine/Nemotron-3-Ultra(Global)

    Proposed, not registered yet

  • InklingComing soon

    Thinking Machines Lab · Global · Chat completions

    A multimodal mixture of experts, 975B total and 41B active, reasoning over text, images and audio at a 256k context.

    Input, Rs / 1M
    Rs 300.78
    Output, Rs / 1M
    Rs 1,218.16

    Route id

    paranine/Inkling(Global)

    Proposed, not registered yet

  • Inkling-SmallComing soon

    Thinking Machines Lab · Global · Chat completions

    The smaller, open-weight Inkling: faster and cheaper to run, with the same native text, image and audio input.

    Input, Rs / 1M
    Rs 150.39
    Output, Rs / 1M
    Rs 360.94

    Route id

    paranine/Inkling-Small(Global)

    Proposed, not registered yet

PK NPU-native endpoints (Ascend 910B)

Served from P/9's own Ascend 910B NPU capacity in Pakistan, the platform's native hardware. Open on the same endpoint as every other route, with the same keys and the same rupee billing.

  • GLM-5.2Live

    Z.ai · PK NPU · Chat completions

    The previous generation of Z.ai's agentic engineering line: coding, tool use and sustained execution on long tasks.

    Input, Rs / 1M
    Rs 421.09
    Output, Rs / 1M
    Rs 1,365.54

    Route id

    paranine/GLM-5.2
  • gpt-oss 120BLive

    OpenAI · PK NPU · Chat completions

    Input, Rs / 1M
    Rs 30.08
    Output, Rs / 1M
    Rs 180.47

    Route id

    paranine/gpt-oss-120b(high)
  • Qwen3.8 27BLive

    Qwen team, Alibaba · PK NPU · Chat completions

    Input, Rs / 1M
    Rs 80.00
    Output, Rs / 1M
    Rs 650.00

    Route id

    paranine/qwen-3.8-27b
  • Qwen3 Embedding 8BLive

    Qwen team, Alibaba · PK NPU · Embeddings

    Input, Rs / 1M
    Rs 2.78
    Output, Rs / 1M
    No output tokens

    Route id

    paranine/qwen3-embedding

How to read this table.

Coming soon means the tag, not the price
A route listed ahead of opening answers 404 until the tag comes off. Its rate is published now so an integration can be costed before it ships, and it is the rate the meter will use on day one.
A proposed id is the shape, not the string
Some routes carry an id P/9 intends to publish but has not registered on the gateway yet. Those say so under the id. Write against them if you like, but expect to change the string once the route opens.
Prices are per million tokens, input and output metered apart
A request is billed on the tokens it actually consumes: the prompt at the input rate, the completion at the output rate. Embedding routes return vectors rather than tokens, so they have no output rate at all.
There is no listing endpoint yet
GET /v1/models is not served, so a stock OpenAI SDK's models.list() returns 404. This page, the dashboard's Model APIs panel and a key's allowed_models are the sources of truth until it ships.

Cost-effective inference for every leading open model.