Authentication
Authentication relies on a single API key passed in the Authorization header. This key is generated immediately after you sign up with an email and password on the Get API key page. You can regenerate the key at any time, which instantly revokes the old one. The API does not require a credit card for the initial trial, and prompts are not used for training your data. Ensure you store your key securely, as it provides direct access to your prepaid credit.
Chat Completions Endpoint
Send requests to the base URL https://api.nsfwllmrouter.com/v1 using the standard POST /v1/chat/completions endpoint. This is a text-only interface; we do not support embeddings, images, audio, or video generation. Specify the model as uncensored to ensure you are using our dedicated open-weight model optimized for adult content. The model runs on our own GPU servers and does not refuse lawful adult, fictional, or controversial topics. It is not GPT, Claude, or any other vendor's model.
curl https://api.nsfwllmrouter.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
The request body must not exceed 8 MB. You can define system, user, and assistant messages just as you would with any OpenAI-compatible client. The model will return text output directly in the response stream.
Python SDK
Use the official OpenAI Python library to integrate quickly. Point the client to our base URL and provide your API key. This approach works for any OpenAI-compatible SDK, allowing you to swap the base URL without changing your application logic. The uncensored model ID ensures you are interacting with our specific uncensored LLM.
from openai import OpenAI
client = OpenAI(base_url="https://api.nsfwllmrouter.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Remember that this is a dedicated uncensored api, so you do not need to manage complex routing or select between multiple models. The system handles the inference on our infrastructure. If you need multi-model routing, you would need to build that yourself, but here the choice is fixed to our single optimized model.
Node SDK
For Node.js environments, initialize the OpenAI client with the custom base URL and your API key. This allows you to use standard JavaScript/TypeScript patterns for building AI applications. The client will communicate with our endpoints exactly as expected by the OpenAI specification.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.nsfwllmrouter.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Ensure your environment supports the required fetch or http libraries. The response object will contain the standard choices array with the generated text. This method is ideal for server-side rendering or API backends where you need reliable, unfiltered text generation.
Streaming Responses
Enable streaming by setting stream: true in your request. The API returns Server-Sent Events (SSE) containing partial chunks of the response. This is useful for real-time interfaces where you want to display text as it is generated. The model supports standard streaming protocols, so you can use existing SSE parsers.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Each chunk contains the incremental text. Aggregate these chunks in your application to construct the full response. Streaming does not affect the token billing; you are charged for the total input and output tokens processed. This feature is supported on the chat/completions endpoint.
Rate Limits and Constraints
Each API key is limited to 300 requests per minute. If you exceed this, you will receive a 429 rate limit error. The maximum request body size is 8 MB. If your key is invalid, you will get a 401 error. If you have no prepaid credit, you will receive a 402 error. The context window is 64,000 tokens, covering both prompt and completion. There is no SLA guarantee, and we do not offer on-prem deployment or model routing. Use this uncensored api for applications that require high volume and unfiltered output.