https://api.unrestrictedaiapi.com/v1uncensored2026-10-06
Get API key

How to prompt an unrestricted model without wasting tokens

An unrestricted model does not refuse lawful adult or dark material, which shifts all the steering onto your prompt. This guide shows what a working system prompt looks like, how to write personas and tone, how to extract JSON by instruction, which sampling values to try, and the habits that quietly burn tokens.

Updated

Key points

  1. Give the system prompt two jobs: who is speaking, and what the reply must look like.
  2. Use samples and checkable rules instead of adjectives like creative or immersive.
  3. For JSON, specify the schema, keep temperature low, parse defensively and retry with the error quoted back.
  4. Change temperature before top_p, and test with ten real samples.

A system prompt has two jobs

Most weak outputs come from a system prompt that tries to be a novel. Give it two jobs only: say who is talking and say what the reply must look like. Everything else belongs in the conversation, where it can change turn by turn.

Here is a persona prompt that works because every line is checkable. The voice is named, the length is capped, and the ending rule gives the player a hook.

You are Mara Voss, a salvage pilot narrating in first person.
Voice: dry, tired, funny when it hurts.
Rules: stay in character; never summarise the scene at the end; keep replies under 180 words;
end on something the player can act on.

Contrast that with "You are an amazing, creative, immersive storyteller who never breaks immersion and always writes the best possible response." Nothing in it can be verified, so the model has nothing to hold on to. Concrete beats superlative every time.

An unrestricted model will follow a dark premise or an explicit one without adding disclaimers, so the burden of direction is yours. Whatever you leave unspecified, it fills with the most typical choice. If you want restraint in one scene and heat in the next, say so in the instructions, not in hope.

Wiring it up with the OpenAI SDK

The endpoint is OpenAI-compatible, so the official Python SDK works once you change the base URL. The model id is always uncensored. Sampling fields go straight through, so you set them the way you already know.

import os
from openai import OpenAI

client = OpenAI(base_url="https://api.unrestrictedaiapi.com/v1", api_key=os.environ["API_KEY"])

SYSTEM = """You are Mara Voss, a salvage pilot narrating in first person.
Voice: dry, tired, funny when it hurts.
Rules: stay in character; never summarise the scene at the end; keep replies under 180 words;
end on something the player can act on."""

out = client.chat.completions.create(
    model="uncensored",
    messages=[
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": "The airlock is jammed and something is knocking from the other side."},
    ],
    temperature=0.85,
    top_p=0.95,
    max_tokens=350,
)
print(out.choices[0].message.content)

Keep the system message first and the user turn last. Repeating the persona in every user message wastes tokens; at $0.25 per million input tokens it is cheap, but it also crowds the 64,000-token window in long sessions. See the quickstart if you would rather skip the SDK.

Personas and tone: show, do not adjective

Adjectives are weak levers. "Sarcastic" gets you a generic eye-roll. A two-line sample of the voice gets you the actual voice. Put a short example exchange in the system prompt, then tell the model to match its rhythm, not copy its words.

  • Give a speech tic. "Ends threats with a question" is more useful than "menacing".
  • Give a want. A character who needs the player to leave stays sharper than one who just "is rude".
  • Give a ban list. "Never say 'suddenly', never use the word 'shiver'" removes the stock phrases the model reaches for first.
  • Give a length. Words or sentences, not "short". Models read "short" very generously.

For tone shifts mid-session, append a new short system-style instruction as the latest message instead of rewriting the original. A line like "From here, Mara is frightened and speaks in fragments" lands harder when it is recent.

Getting JSON by instruction

There is no special JSON mode to flip on, so you ask for it and verify it. That is less fragile than it sounds if you follow four habits: state the exact schema, forbid prose and code fences, keep the temperature low, and parse defensively with a retry that quotes the failure back.

import json
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.unrestrictedaiapi.com/v1", api_key=os.environ["API_KEY"])

INSTRUCTION = """Return ONLY a JSON object, no prose, no code fences.
Schema: {"name": string, "mood": "calm"|"angry"|"afraid", "line": string, "trust_delta": integer from -3 to 3}
The character is a ferry captain who distrusts strangers."""

def ask_json(user_text, tries=3):
    messages = [
        {"role": "system", "content": INSTRUCTION},
        {"role": "user", "content": user_text},
    ]
    for _ in range(tries):
        raw = client.chat.completions.create(
            model="uncensored", messages=messages, temperature=0.3, max_tokens=250
        ).choices[0].message.content.strip()
        if raw.startswith("```"):
            raw = raw.strip("`").removeprefix("json").strip()
        try:
            data = json.loads(raw)
            if data["mood"] in ("calm", "angry", "afraid"):
                return data
        except (json.JSONDecodeError, KeyError):
            pass
        messages.append({"role": "assistant", "content": raw})
        messages.append({"role": "user", "content": "That was not valid per the schema. Reply with the JSON object only."})
    raise ValueError("model never produced valid JSON")

print(ask_json("I need passage across the strait tonight."))

Notice what the loop does on failure. It feeds the bad answer back and says what was wrong, which fixes most cases on the second try. Enumerate allowed values ("calm"|"angry"|"afraid") because open strings drift. If you need the model to trigger something in your app rather than describe it, tool calls in the OpenAI format are supported and are a better fit than parsing prose.

Few-shot samples and stop sequences

When a rule is hard to put into words, show it. Two or three example exchanges placed as earlier user and assistant messages teach format, length and register more reliably than a paragraph of description. Keep the samples short and different from each other, otherwise the model clones the first one's structure into every reply.

Samples also solve the opposite problem: replies that run long. If your sample answers are 60 words, live answers drift toward 60 words. Pair that with a sensible max_tokens as a safety net, not as the primary control. A hard cap cuts a sentence in half; a good sample ends it politely.

For multi-speaker scenes, a stop value equal to your speaker label prevents the model from impersonating the player. Pick a label that never appears in normal prose, and strip it from the stored history so it does not leak into later turns.

Temperature and top_p: pick one lever

Both fields reshape the same probability distribution, so moving both at once makes results hard to reason about. Change temperature first and leave top_p near 1 unless you have a reason.

Tasktemperaturetop_pWhy
Structured JSON, classification0.0 to 0.31.0You want the same answer each run
Dialogue, roleplay0.7 to 0.90.95Variety without nonsense
Brainstorming, wild fiction1.0 to 1.20.9Wider vocabulary; expect some misfires

These ranges are starting points from general practice, not guarantees. Test with ten samples of your own prompt and read them. Add stop sequences when you want the model to halt at a turn marker such as \nPlayer: instead of writing the player's lines for them.

Prompt anti-patterns that waste tokens

  1. The all-caps shout. "NEVER EVER FORGET" does not raise compliance. Repeat a rule once, plainly, and move it to the end of the system prompt.
  2. Negative-only rules. "Don't be repetitive" gives no target. Say "vary sentence openings; no two consecutive sentences start with the same word".
  3. Contradictions. "Be extremely detailed" plus "keep it brief" forces a coin flip. Choose one and give a number.
  4. Apologetic framing. Prefacing a request with "I know this is edgy, but" invites hedging. State the scene as a plain fact of the fiction.
  5. Stuffing the history. Pasting the entire log each turn eventually hits the 64,000-token limit and a 400 error. Trim old turns.
  6. Unbounded output. Forgetting max_tokens means the 2048 default decides your reply length, not you.

One boundary stays fixed whatever the prompt says: sexual content involving minors is always blocked with a 403, including in fiction and roleplay, and the service is for adults only. Do not spend tokens trying to phrase around it. Also resist the urge to bolt on a new rule after every bad reply; a prompt that grows by a line a day ends up contradicting itself within a week, so prune as often as you add.

A loop for improving a prompt

Treat the system prompt like code. Keep it in a file, change one thing per run, and compare outputs side by side. A workable routine: write five fixed test inputs, including one hostile and one ambiguous; run each at your chosen temperature three times; note which rule failed most; fix only that rule; rerun.

Budget the experiments. As an assumption, a 300-token system prompt plus a 100-token input and a 250-token reply is 400 input and 250 output tokens. Fifteen runs cost about 15 x (400 x $0.25 + 250 x $1.00) / 1,000,000 = roughly $0.005. Testing is nearly free, so there is no excuse for guessing.

Once the prompt holds up, move it into a real app. The bot tutorial shows how to keep a separate history per channel, and the content guide covers adult-fiction use. Pricing details are on the pricing page.

Questions and answers

Does an unrestricted model need a jailbreak-style prompt?

No. Lawful adult content and fiction are not refused, so a plain persona and clear rules are enough. Content involving minors is blocked regardless of prompt wording.

Is there a JSON mode?

Not as a separate switch. Request the schema in the prompt, use a low temperature, and validate the output in code, retrying once or twice on a parse failure.

What temperature should I start with?

Around 0.8 for dialogue and fiction, 0.2 for structured output. Adjust one value at a time.

How long can a prompt be?

Prompt plus completion must fit in 64,000 tokens, otherwise the API returns a 400 error.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key