A system prompt has two jobs
Most weak outputs come from a system prompt that tries to be a novel. Give it two jobs only: say who is talking and say what the reply must look like. Everything else belongs in the conversation, where it can change turn by turn.
Here is a persona prompt that works because every line is checkable. The voice is named, the length is capped, and the ending rule gives the player a hook.
You are Mara Voss, a salvage pilot narrating in first person.
Voice: dry, tired, funny when it hurts.
Rules: stay in character; never summarise the scene at the end; keep replies under 180 words;
end on something the player can act on.Contrast that with "You are an amazing, creative, immersive storyteller who never breaks immersion and always writes the best possible response." Nothing in it can be verified, so the model has nothing to hold on to. Concrete beats superlative every time.
An unrestricted model will follow a dark premise or an explicit one without adding disclaimers, so the burden of direction is yours. Whatever you leave unspecified, it fills with the most typical choice. If you want restraint in one scene and heat in the next, say so in the instructions, not in hope.
Wiring it up with the OpenAI SDK
The endpoint is OpenAI-compatible, so the official Python SDK works once you change the base URL. The model id is always uncensored. Sampling fields go straight through, so you set them the way you already know.
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.unrestrictedaiapi.com/v1", api_key=os.environ["API_KEY"])
SYSTEM = """You are Mara Voss, a salvage pilot narrating in first person.
Voice: dry, tired, funny when it hurts.
Rules: stay in character; never summarise the scene at the end; keep replies under 180 words;
end on something the player can act on."""
out = client.chat.completions.create(
model="uncensored",
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": "The airlock is jammed and something is knocking from the other side."},
],
temperature=0.85,
top_p=0.95,
max_tokens=350,
)
print(out.choices[0].message.content)Keep the system message first and the user turn last. Repeating the persona in every user message wastes tokens; at $0.25 per million input tokens it is cheap, but it also crowds the 64,000-token window in long sessions. See the quickstart if you would rather skip the SDK.
Personas and tone: show, do not adjective
Adjectives are weak levers. "Sarcastic" gets you a generic eye-roll. A two-line sample of the voice gets you the actual voice. Put a short example exchange in the system prompt, then tell the model to match its rhythm, not copy its words.
- Give a speech tic. "Ends threats with a question" is more useful than "menacing".
- Give a want. A character who needs the player to leave stays sharper than one who just "is rude".
- Give a ban list. "Never say 'suddenly', never use the word 'shiver'" removes the stock phrases the model reaches for first.
- Give a length. Words or sentences, not "short". Models read "short" very generously.
For tone shifts mid-session, append a new short system-style instruction as the latest message instead of rewriting the original. A line like "From here, Mara is frightened and speaks in fragments" lands harder when it is recent.
Getting JSON by instruction
There is no special JSON mode to flip on, so you ask for it and verify it. That is less fragile than it sounds if you follow four habits: state the exact schema, forbid prose and code fences, keep the temperature low, and parse defensively with a retry that quotes the failure back.
import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.unrestrictedaiapi.com/v1", api_key=os.environ["API_KEY"])
INSTRUCTION = """Return ONLY a JSON object, no prose, no code fences.
Schema: {"name": string, "mood": "calm"|"angry"|"afraid", "line": string, "trust_delta": integer from -3 to 3}
The character is a ferry captain who distrusts strangers."""
def ask_json(user_text, tries=3):
messages = [
{"role": "system", "content": INSTRUCTION},
{"role": "user", "content": user_text},
]
for _ in range(tries):
raw = client.chat.completions.create(
model="uncensored", messages=messages, temperature=0.3, max_tokens=250
).choices[0].message.content.strip()
if raw.startswith("```"):
raw = raw.strip("`").removeprefix("json").strip()
try:
data = json.loads(raw)
if data["mood"] in ("calm", "angry", "afraid"):
return data
except (json.JSONDecodeError, KeyError):
pass
messages.append({"role": "assistant", "content": raw})
messages.append({"role": "user", "content": "That was not valid per the schema. Reply with the JSON object only."})
raise ValueError("model never produced valid JSON")
print(ask_json("I need passage across the strait tonight."))Notice what the loop does on failure. It feeds the bad answer back and says what was wrong, which fixes most cases on the second try. Enumerate allowed values ("calm"|"angry"|"afraid") because open strings drift. If you need the model to trigger something in your app rather than describe it, tool calls in the OpenAI format are supported and are a better fit than parsing prose.
Few-shot samples and stop sequences
When a rule is hard to put into words, show it. Two or three example exchanges placed as earlier user and assistant messages teach format, length and register more reliably than a paragraph of description. Keep the samples short and different from each other, otherwise the model clones the first one's structure into every reply.
Samples also solve the opposite problem: replies that run long. If your sample answers are 60 words, live answers drift toward 60 words. Pair that with a sensible max_tokens as a safety net, not as the primary control. A hard cap cuts a sentence in half; a good sample ends it politely.
For multi-speaker scenes, a stop value equal to your speaker label prevents the model from impersonating the player. Pick a label that never appears in normal prose, and strip it from the stored history so it does not leak into later turns.
Temperature and top_p: pick one lever
Both fields reshape the same probability distribution, so moving both at once makes results hard to reason about. Change temperature first and leave top_p near 1 unless you have a reason.
| Task | temperature | top_p | Why |
|---|---|---|---|
| Structured JSON, classification | 0.0 to 0.3 | 1.0 | You want the same answer each run |
| Dialogue, roleplay | 0.7 to 0.9 | 0.95 | Variety without nonsense |
| Brainstorming, wild fiction | 1.0 to 1.2 | 0.9 | Wider vocabulary; expect some misfires |
These ranges are starting points from general practice, not guarantees. Test with ten samples of your own prompt and read them. Add stop sequences when you want the model to halt at a turn marker such as \nPlayer: instead of writing the player's lines for them.
Prompt anti-patterns that waste tokens
- The all-caps shout. "NEVER EVER FORGET" does not raise compliance. Repeat a rule once, plainly, and move it to the end of the system prompt.
- Negative-only rules. "Don't be repetitive" gives no target. Say "vary sentence openings; no two consecutive sentences start with the same word".
- Contradictions. "Be extremely detailed" plus "keep it brief" forces a coin flip. Choose one and give a number.
- Apologetic framing. Prefacing a request with "I know this is edgy, but" invites hedging. State the scene as a plain fact of the fiction.
- Stuffing the history. Pasting the entire log each turn eventually hits the 64,000-token limit and a 400 error. Trim old turns.
- Unbounded output. Forgetting
max_tokensmeans the 2048 default decides your reply length, not you.
One boundary stays fixed whatever the prompt says: sexual content involving minors is always blocked with a 403, including in fiction and roleplay, and the service is for adults only. Do not spend tokens trying to phrase around it. Also resist the urge to bolt on a new rule after every bad reply; a prompt that grows by a line a day ends up contradicting itself within a week, so prune as often as you add.
A loop for improving a prompt
Treat the system prompt like code. Keep it in a file, change one thing per run, and compare outputs side by side. A workable routine: write five fixed test inputs, including one hostile and one ambiguous; run each at your chosen temperature three times; note which rule failed most; fix only that rule; rerun.
Budget the experiments. As an assumption, a 300-token system prompt plus a 100-token input and a 250-token reply is 400 input and 250 output tokens. Fifteen runs cost about 15 x (400 x $0.25 + 250 x $1.00) / 1,000,000 = roughly $0.005. Testing is nearly free, so there is no excuse for guessing.
Once the prompt holds up, move it into a real app. The bot tutorial shows how to keep a separate history per channel, and the content guide covers adult-fiction use. Pricing details are on the pricing page.