Simple
One API to reach every model we serve. Integrate once, then pick models per use case without touching your integration again.
OpenAI-compatible API
Bali LLM is an AI inference platform. Access leading open-source language models through one simple, reliable and developer-friendly API — no GPU infrastructure, no model deployment, no new SDK to learn.
/v1/chat/completionsWhat is Bali LLM
Bali LLM makes it easy for developers and businesses to integrate large language models into their applications. We host and serve leading open-source models behind a single endpoint, so your team ships features instead of provisioning GPUs.
Because the API follows the OpenAI specification, moving an existing application over is
usually a change of base_url and a model name — the rest of your code stays
exactly as it is.
Why Bali LLM
One API, one key, one bill — and the freedom to change your mind about models.
One API to reach every model we serve. Integrate once, then pick models per use case without touching your integration again.
Low-latency inference tuned for production traffic, with streaming responses so chat and agent workloads feel instant.
Choose the model that fits the job — a small fast one for classification, a larger one for reasoning over long documents.
Scale AI workloads up and down without buying GPUs, managing clusters, or planning capacity months ahead.
Already using OpenAI? Switching is simple. Works with the OpenAI SDKs, LangChain, LlamaIndex, n8n, and anything that speaks the same spec.
Dedicated endpoints, higher rate limits, private deployment, and dedicated capacity when a shared endpoint is not enough.
Available models
Same API, same request format — different cost and capability points.
Powered by leading open-source models
| Model | Type | Context | Use case | Status |
|---|---|---|---|---|
| Kimi K3 | Chat · Reasoning | 1M | Flagship — agents, multi-step reasoning, very long documents | Available |
| Nemotron 3.5 Lightning | Chat | 1M | Long-context chat and high-throughput drafting | Available |
| Qwen 3.5 122B | Chat | 256K | Long documents, summarisation, RAG | Available |
| Qwen 3.6 35B | Chat | 256K | Fast general chat, customer support, classification | Available |
Need a dedicated endpoint or a private deployment? Tell us about your use case — pricing is flexible and quoted per workload.
Quickstart
If you already use an OpenAI-compatible API, getting started takes minutes. Point
base_url at Bali LLM, pick a model, and send your first request — the rest
of your application stays exactly as it is.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_BALILLM_API_KEY",
base_url="https://api.balillm.ai/v1",
)
response = client.chat.completions.create(
model="kimi-k3",
messages=[
{"role": "user", "content": "Hello BaliLLM"},
],
)
print(response.choices[0].message.content)
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.BALILLM_API_KEY,
baseURL: "https://api.balillm.ai/v1",
});
const res = await client.chat.completions.create({
model: "kimi-k3",
messages: [
{ role: "user",
content: "Buatkan balasan email penawaran sewa menara, sopan dan singkat." },
],
});
console.log(res.choices[0].message.content);
curl https://api.balillm.ai/v1/chat/completions \
-H "Authorization: Bearer $BALILLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"messages": [
{ "role": "user", "content": "Apa itu Bali LLM?" }
]
}'
Already using an OpenAI-compatible API? Changing base_url is usually the whole migration.
Use cases
Conversational assistants for customer service, in Bahasa Indonesia or any language your users speak.
Tool-using agents and automation, powered by function calling and structured JSON output.
Generate, translate, and transform content at scale — from product copy to internal reports.
RAG over SOPs, contracts, and regulations, wired into the business systems you already run.
Start integrating powerful AI models with Bali LLM. Tell us what you are building and we will get you a key.
Contact us
Leave your details and a short note about your use case. Our team will reply by email — usually within one business day.
Need a dedicated AI solution?