Get API key

HomeDocs

Llama API: First request in five minutes

Get your first response from an uncensored LLM in under five minutes. This quickstart shows how to replace your existing OpenAI client configuration with our base URL and API key.

https://api.llamaapis.com/v1uncensored

llamaapis.com

Installation

Start by installing the official OpenAI Python client. It handles the HTTP requests and JSON parsing for you.

  • pip install openai

You can also use the JavaScript SDK for Node.js. Both libraries are fully compatible with our API because we follow the OpenAI chat-completions specification exactly.

Authentication Setup

Our API uses standard API key authentication. After signing up at https://llamaapis.com, you receive a key immediately. No credit card is required for the trial tier.

Store this key in an environment variable or pass it directly to your client. The key authenticates all requests to the /v1 endpoints. If you lose your key, you can regenerate it from your dashboard, which invalidates the old one immediately.

Basic Chat Completion Request

Make a request to the chat completions endpoint. You must set the base URL to our server and provide your API key.

The model ID is uncensored. This model does not apply standard content filters for lawful adult use.

Here is how you make a simple request using curl:

curl https://api.llamaapis.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

The response returns a generated text string. If you need function calling, the structure is identical to OpenAI's format.

Streaming Responses (SSE)

For real-time output, enable streaming by setting stream: true. The server sends Server-Sent Events (SSE) containing partial deltas.

This is useful for chat interfaces where you want to show text as it is generated. The context window supports up to 100,000 tokens, so streaming works efficiently even for long conversations.

Example streaming implementation:

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Remember to handle the stop event to close the connection cleanly.

Tool/Function Calling Support

Our API supports structured tool calling. Define your functions in the tools parameter and set the model to respond with function calls.

The client library will parse the JSON arguments automatically. This works exactly as documented in the OpenAI specification.

Example using the Python SDK:

from openai import OpenAI

client = OpenAI(base_url="https://api.llamaapis.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

The model will return the function name and arguments. You execute the function and send the result back in the next message.

Limits, Errors, and Context

Be aware of the following limits:

  • Rate limit: 300 requests per minute per API key.
  • Max request body: 8 MB.
  • Context window: 100,000 tokens total (input + output).

Common errors:

  • 401 Unauthorized: Invalid or missing API key.
  • 402 Payment Required: Insufficient prepaid credit.
  • 429 Too Many Requests: You exceeded the 300 rpm limit.

We do not offer embeddings, image generation, or audio processing. This is a text-only LLM API.

Node.js

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.llamaapis.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

What the API supports

A quick checklist for developers: format, limits, features, billing.

ItemValue
CompatibilityOpenAI Chat Completions schema; official openai SDKs work unchanged
AuthenticationAuthorization: Bearer YOUR_KEY
MethodsPOST /v1/chat/completions · GET /v1/models
Base URLhttps://api.llamaapis.com/v1
Modeluncensored
Max context100,000 tokens, input and output combined
Completion lengthup to 16,000 tokens per request (default 2,048)
Tools / tool callsYes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool
Structured outputJSON object mode via response_format json_object
Sampling parameterstemperature, top_p, stop, seed and the two penalties are passed through
SSE streamingSupported (stream: true), usage included at the end
Requests per minute300 requests per minute per key
Concurrency8 requests at the same time per key
Request size8 MB request body
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Volume bonus+5% from $50, +10% from $100
Token prices$0.25 per 1M input tokens · $1.00 per 1M output tokens
Billingprepaid credit, charged by real token usage; errors and refusals are free
Top-upUSDT (TRC20) or USDC (Base), any whole amount from $10 to $500
Subscriptionpaid credit never expires, no subscription
Free trial$0.50 for 7 days, no card
Keysone key per account, regenerate any time (the old one stops working)
Content policyuncensored for adults; the only hard rule: no sexual content involving minors
Accountsign in with Google or with e-mail + password

Error codes

The type field is stable, the message is for humans. Errors cost nothing.

CodeTypeMeaning
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditbalance is empty — top up, requests resume at once
403content_blockedsexual content involving minors — refused, not billed
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largebody over 8 MB
429rate_limited · concurrencyslow down: rate or parallel limit reached
503upstream_busytemporary overload, retry shortly

Questions and answers

llamaapis.com
What happens if I exceed my prepaid credit?

Your requests will receive a 402 error until you top up your account. You can add credit via crypto (USDT or USDC) starting at $10. Existing prepaid credit never expires.

Is the uncensored model fine-tuned or base?

It is an open-weight model run on our own GPU servers, specifically tuned to answer without content refusals for lawful adult use. It is not a direct copy of GPT or Claude.

Can I use the same SDK code I use with OpenAI?

Yes. Change the <code>base_url</code> to <code>https://api.llamaapis.com/v1</code> and update the API key. The request and response formats are compatible.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key