Home/Docs
Llama API: First request in five minutes
Get your first response from an uncensored LLM in under five minutes. This quickstart shows how to replace your existing OpenAI client configuration with our base URL and API key.
https://api.llamaapis.com/v1uncensored
Installation
Start by installing the official OpenAI Python client. It handles the HTTP requests and JSON parsing for you.
- pip install openai
You can also use the JavaScript SDK for Node.js. Both libraries are fully compatible with our API because we follow the OpenAI chat-completions specification exactly.
Authentication Setup
Our API uses standard API key authentication. After signing up at https://llamaapis.com, you receive a key immediately. No credit card is required for the trial tier.
Store this key in an environment variable or pass it directly to your client. The key authenticates all requests to the /v1 endpoints. If you lose your key, you can regenerate it from your dashboard, which invalidates the old one immediately.
Basic Chat Completion Request
Make a request to the chat completions endpoint. You must set the base URL to our server and provide your API key.
The model ID is uncensored. This model does not apply standard content filters for lawful adult use.
Here is how you make a simple request using curl:
curl https://api.llamaapis.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'The response returns a generated text string. If you need function calling, the structure is identical to OpenAI's format.
Streaming Responses (SSE)
For real-time output, enable streaming by setting stream: true. The server sends Server-Sent Events (SSE) containing partial deltas.
This is useful for chat interfaces where you want to show text as it is generated. The context window supports up to 100,000 tokens, so streaming works efficiently even for long conversations.
Example streaming implementation:
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Remember to handle the stop event to close the connection cleanly.
Tool/Function Calling Support
Our API supports structured tool calling. Define your functions in the tools parameter and set the model to respond with function calls.
The client library will parse the JSON arguments automatically. This works exactly as documented in the OpenAI specification.
Example using the Python SDK:
from openai import OpenAI
client = OpenAI(base_url="https://api.llamaapis.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)The model will return the function name and arguments. You execute the function and send the result back in the next message.
Limits, Errors, and Context
Be aware of the following limits:
- Rate limit: 300 requests per minute per API key.
- Max request body: 8 MB.
- Context window: 100,000 tokens total (input + output).
Common errors:
- 401 Unauthorized: Invalid or missing API key.
- 402 Payment Required: Insufficient prepaid credit.
- 429 Too Many Requests: You exceeded the 300 rpm limit.
We do not offer embeddings, image generation, or audio processing. This is a text-only LLM API.
Node.js
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.llamaapis.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);What the API supports
A quick checklist for developers: format, limits, features, billing.
| Item | Value |
|---|---|
| Compatibility | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Authentication | Authorization: Bearer YOUR_KEY |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Base URL | https://api.llamaapis.com/v1 |
| Model | uncensored |
| Max context | 100,000 tokens, input and output combined |
| Completion length | up to 16,000 tokens per request (default 2,048) |
| Tools / tool calls | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Structured output | JSON object mode via response_format json_object |
| Sampling parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| SSE streaming | Supported (stream: true), usage included at the end |
| Requests per minute | 300 requests per minute per key |
| Concurrency | 8 requests at the same time per key |
| Request size | 8 MB request body |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Volume bonus | +5% from $50, +10% from $100 |
| Token prices | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Top-up | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Subscription | paid credit never expires, no subscription |
| Free trial | $0.50 for 7 days, no card |
| Keys | one key per account, regenerate any time (the old one stops working) |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
| Account | sign in with Google or with e-mail + password |
Error codes
The type field is stable, the message is for humans. Errors cost nothing.
| Code | Type | Meaning |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
What happens if I exceed my prepaid credit?
Your requests will receive a 402 error until you top up your account. You can add credit via crypto (USDT or USDC) starting at $10. Existing prepaid credit never expires.
Is the uncensored model fine-tuned or base?
It is an open-weight model run on our own GPU servers, specifically tuned to answer without content refusals for lawful adult use. It is not a direct copy of GPT or Claude.
Can I use the same SDK code I use with OpenAI?
Yes. Change the <code>base_url</code> to <code>https://api.llamaapis.com/v1</code> and update the API key. The request and response formats are compatible.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.