https://api.unrestrictedaiapi.com/v1uncensored2026-10-06
Get API key

Build a Discord or Telegram bot on an unrestricted AI API

A chat bot is the fastest way to put an unrestricted model in front of people. This walkthrough builds both a discord.py bot and a python-telegram-bot bot on one shared module that handles the API call, so history per channel, rate limiting, retries and an age-gate are written once and reused.

Updated

Key points

  1. Isolate API logic in one module and keep each platform adapter tiny.
  2. Keep a separate bounded history per channel or chat so conversations never mix.
  3. Throttle in three layers: per-user cooldown, a global sliding window under 300 requests a minute, and a concurrency cap.
  4. Gate adult use: NSFW-flagged Discord channels only, and explicit /adult confirmation on Telegram.

What we are building

Two small bots, one shared brain. A file called llm.py owns everything that touches the API: the persona, the rate limiter, the retry logic. A Discord bot and a Telegram bot then do nothing except translate platform events into a list of messages and back. Add a third platform later and you only write the adapter.

Install the dependencies with pip install httpx discord.py python-telegram-bot. Export API_KEY from your account page, plus DISCORD_TOKEN and/or TELEGRAM_TOKEN from each platform's own bot setup. Both bots below assume Python 3.10 or newer.

The request format is the plain OpenAI chat-completions shape, covered in the multi-language quickstart. Nothing here needs an SDK.

The shared core: persona, limiter and retries

Put the dangerous part in one place. The module below keeps a sliding window of request timestamps and waits when it nears 240 calls a minute, a deliberate margin under the 300 per minute allowed per key. A semaphore caps concurrency at four so a busy channel cannot fire twenty requests at once.

# llm.py - shared by both bots
import asyncio
import os
import time
from collections import deque

import httpx

URL = "https://api.unrestrictedaiapi.com/v1/chat/completions"
PERSONA = "You are Pip, a sardonic bar-room storyteller. Keep replies under 150 words."

_stamps = deque()            # timestamps of recent calls
_gate = asyncio.Semaphore(4) # at most 4 requests in flight
_client = httpx.AsyncClient(timeout=60.0)

async def _wait_for_slot(limit=240, window=60.0):
    # Stay below the 300 requests/minute per-key limit with headroom.
    while True:
        now = time.monotonic()
        while _stamps and now - _stamps[0] > window:
            _stamps.popleft()
        if len(_stamps) < limit:
            _stamps.append(now)
            return
        await asyncio.sleep(window - (now - _stamps[0]) + 0.05)

async def chat(history):
    """history: list of {"role","content"} dicts, oldest first."""
    messages = [{"role": "system", "content": PERSONA}] + list(history)
    async with _gate:
        for attempt in range(3):
            await _wait_for_slot()
            r = await _client.post(
                URL,
                headers={"Authorization": f"Bearer {os.environ['API_KEY']}"},
                json={"model": "uncensored", "messages": messages, "max_tokens": 400},
            )
            if r.status_code in (429, 503):
                await asyncio.sleep(2 ** attempt)
                continue
            if r.status_code == 403:
                return "I can't continue with that."
            r.raise_for_status()
            return r.json()["choices"][0]["message"]["content"]
    return "The service is busy. Try again in a moment."

Look at the status handling. A 429 or 503 sleeps for 1, 2, then 4 seconds and tries again. A 403 means the content filter fired, so the bot answers with a neutral line rather than retrying. Any other error raises, which your platform library will log. If you see a 402, your balance is empty or the trial ended; top up the prepaid balance.

The Discord adapter

Discord bots need the message-content intent switched on in the developer portal as well as in code, or msg.content arrives empty. The bot only answers when mentioned, which keeps it from reading every message in a busy server.

# discord_bot.py
import os
import time
from collections import defaultdict, deque

import discord

from llm import chat

intents = discord.Intents.default()
intents.message_content = True      # also enable it in the developer portal
bot = discord.Client(intents=intents)

HISTORY_TURNS = 10                   # user+assistant messages kept per channel
history = defaultdict(lambda: deque(maxlen=HISTORY_TURNS))
last_use = {}                        # user id -> last request time
COOLDOWN = 5.0                       # seconds between requests per user

def chunks(text, size=1900):
    return [text[i:i + size] for i in range(0, len(text), size)] or [""]

@bot.event
async def on_message(msg: discord.Message):
    if msg.author.bot or bot.user not in msg.mentions:
        return
    # Age gate: guild text channels must be flagged NSFW. No DMs.
    if not isinstance(msg.channel, discord.TextChannel) or not msg.channel.is_nsfw():
        await msg.reply("I only chat in channels marked age-restricted (NSFW).")
        return
    now = time.monotonic()
    if now - last_use.get(msg.author.id, 0) < COOLDOWN:
        await msg.add_reaction("\u23f3")
        return
    last_use[msg.author.id] = now

    text = msg.clean_content.replace(f"@{bot.user.display_name}", "").strip()
    if not text:
        return
    h = history[msg.channel.id]
    h.append({"role": "user", "content": f"{msg.author.display_name}: {text}"})
    async with msg.channel.typing():
        answer = await chat(h)
    h.append({"role": "assistant", "content": answer})
    for part in chunks(answer):
        await msg.channel.send(part)

bot.run(os.environ["DISCORD_TOKEN"])

History is a deque with maxlen=10 keyed by channel id, so each channel is its own conversation and old turns fall off automatically. Prefixing each user line with the speaker's display name lets the model tell people apart in a shared room. Ten messages is a conservative default; the context window is 64,000 tokens, so you can raise it a lot before it matters.

The Telegram adapter

python-telegram-bot v20 and later is async, which pairs naturally with the shared core. chat_data is a per-chat dictionary the library keeps for you, so history and cooldown state need no extra storage code.

# telegram_bot.py  (python-telegram-bot v20+)
import os
import time
from collections import deque

from telegram import Update
from telegram.ext import (ApplicationBuilder, CommandHandler, ContextTypes,
                          MessageHandler, filters)

from llm import chat

async def adult(update: Update, context: ContextTypes.DEFAULT_TYPE):
    context.chat_data["adult"] = True
    await update.message.reply_text("Noted. Say something.")

async def start(update: Update, context: ContextTypes.DEFAULT_TYPE):
    await update.message.reply_text(
        "This bot produces adult fiction. Send /adult only if you are 18 or older."
    )

async def talk(update: Update, context: ContextTypes.DEFAULT_TYPE):
    cd = context.chat_data
    if not cd.get("adult"):
        await update.message.reply_text("Send /adult to confirm you are 18+.")
        return
    if time.monotonic() - cd.get("last", 0) < 4:
        return                                   # per-chat cooldown
    cd["last"] = time.monotonic()

    h = cd.setdefault("history", deque(maxlen=10))
    h.append({"role": "user", "content": update.message.text})
    await context.bot.send_chat_action(update.effective_chat.id, "typing")
    answer = await chat(h)
    h.append({"role": "assistant", "content": answer})
    for i in range(0, len(answer), 4000):        # Telegram limit is 4096 chars
        await update.message.reply_text(answer[i:i + 4000])

app = ApplicationBuilder().token(os.environ["TELEGRAM_TOKEN"]).build()
app.add_handler(CommandHandler("start", start))
app.add_handler(CommandHandler("adult", adult))
app.add_handler(MessageHandler(filters.TEXT & ~filters.COMMAND, talk))
app.run_polling()

run_polling() is the simplest way to start and needs no public URL. Keep in mind that chat_data lives in memory, so a restart clears it. If you need history to survive restarts, store the deque contents yourself in a file or database.

Age-gating and the NSFW channel rule

This API is for adults only, and anything you build on it inherits that responsibility. Treat age-gating as a feature, not an afterthought.

  • Discord: answer only in guild channels flagged age-restricted. The check channel.is_nsfw() costs nothing and the bot refuses elsewhere, including direct messages.
  • Telegram: require an explicit /adult confirmation per chat before any generation, and state the rule in the welcome text.
  • Both: do not offer the bot in spaces aimed at young people, and remove access promptly if you learn a user is under 18.

A confirmation command is a gate, not proof of age. Sexual content involving minors is always blocked by the API with a 403, including fiction and roleplay, and your bot should never try to work around that response.

Running it and testing without a crowd

Start each bot in its own terminal with python discord_bot.py or python telegram_bot.py. For the first run, create a private test server or a chat with only yourself. Walk through a short checklist and you will catch nearly every beginner bug.

  1. Mention the bot in a non-NSFW channel and confirm it refuses. This proves the gate works.
  2. Mention it in a flagged channel and send two messages quickly. The second should earn the hourglass reaction, which proves the cooldown works.
  3. Hold a four-message conversation, then refer back to the first message. If the bot remembers, history is wired correctly.
  4. Temporarily set the API key to a wrong value and confirm the failure shows up in your logs rather than disappearing silently.
  5. Paste a long block of text and watch the reply split into several messages under the platform's length limit.

If replies feel generic, the persona is usually the problem rather than the code. Change one sentence of PERSONA, restart, and compare. Because the persona lives in the shared module, one edit changes both bots at once, which is exactly why the structure is worth the extra file.

When you are ready to deploy, any always-on process manager will do. The bots hold a long-lived connection to their platform and make short outbound HTTPS calls to the API, so they need no inbound ports and almost no memory. Add a restart policy so a crash does not leave your community talking to a silent bot.

Rate limits, history size and cost

Three knobs control how the bot behaves under load. The per-user cooldown stops one person from monopolising the key. The global window protects the 300 requests per minute ceiling. The semaphore smooths bursts. If you run several bot processes on one key, they share that ceiling, so lower each process's limit accordingly.

Estimate spend before you invite a crowd. Assumptions: each call sends about 1,200 input tokens (persona plus ten turns) and returns 250 tokens. One call costs 1,200 x $0.25 / 1,000,000 + 250 x $1.00 / 1,000,000 = $0.0003 + $0.00025 = $0.00055. A server producing 2,000 replies a day would spend about $1.10 a day under those assumptions. Your numbers will differ, so log the usage field for a week and recalculate.

To tune the persona itself, read the prompting guide. Prices are listed on the pricing page.

Hardening before you share the bot

Keep tokens out of source control; read them from the environment as shown. Cap message length before sending it to the API so one huge paste cannot burn your balance. Log errors but not message text, since your users would reasonably expect their chats to stay private. Add a /reset command that clears a channel's deque, which is the quickest fix when a conversation derails. Finally, put a visible note in the bot's profile saying it writes adult fiction and that its replies are generated.

One more design choice deserves a sentence: the bots wait for the full reply instead of streaming it. Editing a chat message token by token runs straight into platform edit limits, and for replies capped near 150 words the wait is short. If you later want live typing in your own web front end, switch on stream: true there, where you control the transport. Until then, the typing indicator you already send is the cheapest way to show the bot is working, and it keeps the code short enough to read in one sitting and audit before you invite anyone in.

Questions and answers

Can one API key serve several bots?

Yes, but they share the 300 requests per minute limit for that key, so split your internal throttle across the processes.

Does the bot remember conversations after a restart?

Not in these examples; history is held in memory. Persist the deque contents to a file or database if you need continuity.

How do I keep minors out?

Restrict Discord replies to age-restricted channels, require a Telegram confirmation command, and avoid youth-oriented spaces. These are gates, not proof, so act on any report quickly.

What happens when the balance runs out?

The API returns 402 with the code no_credit. Catch it in the core module and tell your users the bot is paused.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key