EN ▾

OpenAI-compatible API without filters for adult use

https://api.llmsemcensura.com/v1uncensored

Quick Integration with Our API

Integrate our OpenAI-compatible API in minutes to generate text without adult content refusals. Use the chat-completions endpoint with the 'uncensored' model and have full control over the 100,000-token context.

https://api.llmsemcensura.com/v1uncensored

Authentication and Base URL

Para começar, acesse nossa plataforma e crie uma conta apenas com e-mail e senha. Após a confirmação, sua chave de API será exibida imediatamente. Não é necessário cartão de crédito para o trial. Utilize a base URL https://api.llmsemcensura.com/v1 em seus clientes e defina a chave no cabeçalho Authorization: Bearer SUA_CHAVE. O modelo a ser invocado é identificado por uncensored.

Chat Completions Endpoint

The POST /v1/chat/completions endpoint is the only route needed for text usage. Send messages in the messages structure and receive direct responses. We do not offer embeddings or image generation. The model is open and runs on our GPU servers, without using data to train other models. Set the model to uncensored in requests.

Basic request example:

curl https://api.llmsemcensura.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Streaming Support (SSE)

For applications requiring low perceived latency, enable streaming by setting stream: true. The API will return a Server-Sent Events (SSE) stream, allowing text to be displayed token by token. This is ideal for responsive chat interfaces. The client must process the stream until the done event.

Streaming usage example:

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Function Calling and Tools

The model supports native function calling. You can define tools in the request body so the AI returns structured arguments instead of free text. This facilitates integration with external systems. The response format follows the OpenAI standard, returning tool_calls when appropriate.

Token Management and Limits

The total context (prompt + completion) is limited to 100,000 tokens. The price is $0.25 per 1M input tokens and $1.00 per 1M output tokens. Paid credit never expires. You can top up from $10, with a +5% bonus above $50 and +10% above $100. The initial trial offers $0.50 valid for 7 days.

Python Code Example

The official openai library works natively. Just change the base_url and provide your key. The code below demonstrates a simple interaction. Make sure to install the package via pip before running.

from openai import OpenAI

client = OpenAI(base_url="https://api.llmsemcensura.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node.js

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.llmsemcensura.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Technical specifications

Everything the endpoint can and cannot do, in one place — check it before you top up.

FeatureSupport
CompatibilityOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
AuthenticationAuthorization: Bearer YOUR_KEY
EndpointsPOST /v1/chat/completions · GET /v1/models
Base URLhttps://api.llmsemcensura.com/v1
Model IDuncensored
Sampling parameterstemperature, top_p, stop, seed and the two penalties are passed through
Max outputup to the rest of the 64,000-token window; max_tokens optional (no separate cap)
Structured outputJSON object mode via response_format json_object
StreamingYes — server-sent events; the last chunk carries token usage
Tools / tool callsYes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool
Max context64,000 tokens (prompt + completion together)
Rate limit300/min per key
Max bodyup to 8 MB per request
HeadersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Parallel requests8 requests at the same time per key
Credit expirypaid credit never expires, no subscription
Bonus credit+5% on $50+, +10% on $100+
Billingpay as you go from prepaid credit; nothing is charged for failed or refused requests
Trial credit$0.50 for 7 days, no card · Trial key: 2 parallel requests, 60 req/min; full limits (8 and 300) after first top-up
Price$0.25 per 1M input tokens · $1.00 per 1M output tokens
Paymentcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Key managementone key per account, regenerate any time (the old one stops working)
Content policyuncensored for adults; the only hard rule: no sexual content involving minors
Sign-inGoogle or e-mail and password

Errors and what to do

Every error is JSON with a type you can switch on. You are never charged for an error.

HTTPTypeWhat to do
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditbalance is empty — top up, requests resume at once
403content_blockedrefused by the content policy
404not_foundunknown endpoint
413request_too_largebody over 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busymodel busy — retry in a few seconds

Frequently asked questions

What is the difference between this model and GPT or Claude?

Our model is open, runs on our own hardware, and does not use your prompts to train other models. It is specifically configured to not refuse legal adult content, unlike the standard models from OpenAI or Anthropic. It is not a version of GPT or Claude.

What happens if I exceed the rate limit?

If you exceed 300 requests per minute, you will receive an HTTP 429 error (Too Many Requests). The response body will indicate the exceeded limit. Reduce the call frequency or wait for the 60-second window reset.

How does content blocking work?

The model is 'uncensored' for adult, political and controversial content, but maintains a strict limit: sexual content involving minors is always blocked. We do not encourage illegal use, but we do not apply standard moral filters for adult users.

Your key is one form away

Create an account, copy the key and change the base URL. That's all the configuration.

Get API keyhttps://api.llmsemcensura.com/v1