39 pointsby fionaattelnyx7 hours ago6 comments
  • fionaattelnyx7 hours ago
    Moonshot AI released open weights for Kimi K3 today and it's live on Telnyx Inference. The architecture and Moonshot's own benchmarks are in their technical blog. This post is about running it on Telnyx.

    What we are adding: K3 is now available via the Telnyx Inference API, hosted on GPUs that we own and operate.

    This matters due to the size of this model. A 2.8T model needs dedicated infra to serve well. Because we own and operate the GPUs, we can contorl throughput. with no inter-provider hops and no cloud tenant, latency is minimized. We have GPUs in each of the US, EU, APAC, and MENA. Inference runs in the region you pick, with zero data retention. We do not store prompts or completions after the response returns. Because we own the infra, the per-token price reflects the cost of running the model, not the cost of renting someone else's plus their margin.

    We have not benchmarked K3 ourselves yet. Moonshot's own numbers and early third-party evaluations put it at frontier level for coding and agentic work, trailing only Claude Fable 5 and GPT 5.6 Sol on aggregate. Full breakdown in their blog.

    Pricing on Telnyx: $2.70/1M input tokens, $13.50/1M output tokens, $0.27/1M cached input tokens. Prompt caching enabled by default. Served via an OpenAI-compatible endpoint so you can test easily.

    Model ID: [MODEL_ID] API: https://api.telnyx.com/v2/ai/chat/completions Docs: https://developers.telnyx.com/docs/inference Technical blog (Moonshot): https://www.kimi.com/blog/kimi-k3

    • smallerize7 hours ago
      Very cool. What are your throughput and latency like?
    • marsven_422an hour ago
      [dead]
    • teravor3 hours ago

          > Because we own the infra, the per-token price reflects the cost of running the model, not the cost of renting someone else's plus their margin.
      
      
      is not compatible with

          > Pricing on Telnyx: $2.70/1M input tokens, $13.50/1M output tokens, $0.27/1M cached input tokens.
      
      since you asserted something false and bizarre, how about telling us what is your markup?
      • alexeldeib3 hours ago
        Why are those incompatible? Pricing is 10% under moonshot, and pure infra providers also want money
      • gruez3 hours ago
        I don't get it, what's the contradiction supposed to be?
      • rubslopes3 hours ago
        They are selling it even cheaper than Moonshot AI. Why are you so sure they are lying?
  • theredsix3 hours ago
    10% cheaper than official! Let the inference pricing wars begin!
  • jakswa4 hours ago
    what quantization? FP4?
    • NitpickLawyeran hour ago
      The model is native mxfp4 w/ mxfp8 activations via QAT.
  • buffer_overlord7 hours ago
    That’s huge
  • LoganDark5 hours ago
    Telnyx was cool until they started demanding KYC. I would use them for burner phone numbers until they started saying they needed my government ID. Fuck that.

    OVH same thing. Tried to buy a VPS from them some years back and they said no VPS unless I provided ID. Would not refund me. Tried to dispute but my bank just gave me a credit instead.

    • nujabean hour ago
      There were new regulations passed to combat robocalls that forced these companies to tighten up.
  • hiherer4 hours ago
    [dead]