{ "cells": [ { "cell_type": "markdown", "metadata": {}, "source": [ "# NoTokenLM-Gen-3.6 — Interactive Playground\n", "\n", "A 13.9-million-parameter, byte-level, tokenizer-free language model trained on an English + French text mix. No subword vocabulary, no BPE — just raw UTF-8 bytes in, raw UTF-8 bytes out.\n", "\n", "This notebook lets you:\n", "- Load **NoTokenLM-Gen-3.6** directly from the Hugging Face Hub\n", "- Generate text from your own prompt\n", "- Control **temperature**, **top-k**, and **max new tokens** with an interactive panel\n", "- Re-run generation as many times as you like without reloading the model\n", "\n", "Works out of the box on **Google Colab**, **Kaggle Notebooks**, and any standard Jupyter environment (CPU or GPU).\n", "\n", "> This model is a small text generator (see the [model card](https://huggingface.co/omurberaisik/NoTokenLM-Gen-3.6) for what it's good and not good at). It writes fluent English, but it has no usable world knowledge and is not meant for factual Q&A, math, code, or conversation. French output is unreliable, and prompts in other languages are not reliably continued in the prompt's language.\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 1. Setup\n", "\n", "Installs the required libraries. Safe to re-run — already-satisfied packages are skipped.\n", "\n", "Note: this intentionally does **not** force-upgrade `torch`. Colab and Kaggle ship a pre-matched `torch`/`torchvision` pair; upgrading only `torch` (without also upgrading `torchvision` to a matching build) causes a `torchvision::nms does not exist` / `ModuleNotFoundError` error when `transformers` loads. If you hit that error, the fix is `Runtime → Restart session` and re-running from the top *without* upgrading torch yourself.\n" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "%pip install -q -U \"transformers>=4.44.0\" ipywidgets safetensors huggingface_hub\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 2. Load the model\n", "\n", "Set `REPO_ID` if you're pointing at a different repo (e.g. a fork or a local path). By default this pulls **NoTokenLM-Gen-3.6** from the Hugging Face Hub.\n", "\n", "`trust_remote_code=True` is required — this model ships its own architecture code (RoPE, RMSNorm, SwiGLU, byte-level I/O) alongside the weights.\n" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "import torch\n", "from transformers import AutoModelForCausalLM, AutoTokenizer\n", "\n", "REPO_ID = \"omurberaisik/NoTokenLM-Gen-3.6\" # change this if you're loading a different repo/fork\n", "\n", "device = \"cuda\" if torch.cuda.is_available() else \"cpu\"\n", "print(f\"Using device: {device}\")\n", "\n", "print(f\"Loading model and tokenizer from '{REPO_ID}' ...\")\n", "model = AutoModelForCausalLM.from_pretrained(REPO_ID, trust_remote_code=True)\n", "tokenizer = AutoTokenizer.from_pretrained(REPO_ID, trust_remote_code=True)\n", "\n", "model.to(device)\n", "model.eval()\n", "\n", "n_params = sum(p.numel() for p in model.parameters())\n", "print(f\"Model loaded: {n_params:,} parameters on {device}.\")\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 3. Generation function\n", "\n", "A thin wrapper around the model's own `generate()` method. Byte-level models decode with `errors=\"replace\"` since a generation can legitimately end mid multi-byte UTF-8 character. Byte `0x00` is the model's *end-of-document* marker: by default generation stops when the model emits it, and the marker is dropped from the returned text.\n" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "def run_generation(prompt: str, max_new_tokens: int, temperature: float, top_k: int, do_sample: bool = True, seed: int | None = None):\n", " \"\"\"Generate a continuation for `prompt` and return the full text (prompt + completion).\"\"\"\n", " if seed is not None:\n", " torch.manual_seed(seed)\n", "\n", " if not prompt:\n", " prompt = \" \" # avoid feeding a fully empty byte sequence\n", "\n", " input_ids = torch.tensor([list(prompt.encode(\"utf-8\"))], dtype=torch.long, device=device)\n", "\n", " output_ids = model.generate(\n", " input_ids,\n", " max_new_tokens=max_new_tokens,\n", " temperature=temperature,\n", " top_k=top_k if top_k > 0 else None,\n", " do_sample=do_sample,\n", " )\n", "\n", " byte_values = [b for b in output_ids[0].tolist() if 0 < b <= 255] # 0x00 = end-of-document marker\n", " return bytes(byte_values).decode(\"utf-8\", errors=\"replace\")\n", "\n", "\n", "# Quick sanity check\n", "print(run_generation(\"Once upon a time,\", max_new_tokens=40, temperature=0.5, top_k=40))\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 4. Interactive generation panel\n", "\n", "Adjust the prompt, temperature, top-k, and length below, then click **Generate**. Re-click any time — the model stays loaded in memory, so each generation is fast.\n", "\n", "- **Temperature** — lower (e.g. 0.3) is safer and more repetitive; higher (e.g. 1.0+) is more varied but more likely to drift into incoherence. The model card's own evaluation was run at **temperature 0.5**.\n", "- **Top-k** — restricts sampling to the k most likely next bytes at each step. Lower is safer, higher is more diverse.\n", "- **Max new tokens** — how many new bytes to generate (note: bytes, not characters — most multi-byte UTF-8 characters take 1–4 bytes).\n", "- **Greedy (no sampling)** — when checked, always picks the single most likely next byte (temperature/top-k are ignored). Deterministic, but often more repetitive.\n", "- **Seed** — set a fixed integer for reproducible output, or leave at -1 for a different result every time.\n" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "import ipywidgets as widgets\n", "from IPython.display import display, clear_output\n", "\n", "prompt_box = widgets.Textarea(\n", " value=\"Once upon a time, a little fox\",\n", " placeholder=\"Type a short story opener...\",\n", " description=\"Prompt:\",\n", " layout=widgets.Layout(width=\"100%\", height=\"70px\"),\n", " style={\"description_width\": \"100px\"},\n", ")\n", "\n", "temperature_slider = widgets.FloatSlider(\n", " value=0.5, min=0.05, max=1.5, step=0.05,\n", " description=\"Temperature:\",\n", " style={\"description_width\": \"100px\"},\n", " layout=widgets.Layout(width=\"400px\"),\n", " readout_format=\".2f\",\n", ")\n", "\n", "top_k_slider = widgets.IntSlider(\n", " value=40, min=1, max=200, step=1,\n", " description=\"Top-k:\",\n", " style={\"description_width\": \"100px\"},\n", " layout=widgets.Layout(width=\"400px\"),\n", ")\n", "\n", "max_tokens_slider = widgets.IntSlider(\n", " value=60, min=5, max=512, step=5,\n", " description=\"Max new tokens:\",\n", " style={\"description_width\": \"100px\"},\n", " layout=widgets.Layout(width=\"400px\"),\n", ")\n", "\n", "greedy_checkbox = widgets.Checkbox(\n", " value=False,\n", " description=\"Greedy (no sampling)\",\n", " style={\"description_width\": \"initial\"},\n", ")\n", "\n", "seed_box = widgets.IntText(\n", " value=-1,\n", " description=\"Seed (-1 = random):\",\n", " style={\"description_width\": \"140px\"},\n", " layout=widgets.Layout(width=\"260px\"),\n", ")\n", "\n", "generate_button = widgets.Button(\n", " description=\"Generate\",\n", " button_style=\"primary\",\n", " icon=\"magic\",\n", ")\n", "\n", "output_area = widgets.Output()\n", "\n", "def on_generate_clicked(_):\n", " with output_area:\n", " clear_output(wait=True)\n", " seed = None if seed_box.value < 0 else seed_box.value\n", " print(\"Generating...\\n\")\n", " text = run_generation(\n", " prompt=prompt_box.value,\n", " max_new_tokens=max_tokens_slider.value,\n", " temperature=temperature_slider.value,\n", " top_k=top_k_slider.value,\n", " do_sample=not greedy_checkbox.value,\n", " seed=seed,\n", " )\n", " print(text)\n", "\n", "generate_button.on_click(on_generate_clicked)\n", "\n", "controls = widgets.VBox([\n", " prompt_box,\n", " widgets.HBox([temperature_slider, top_k_slider]),\n", " widgets.HBox([max_tokens_slider, seed_box]),\n", " greedy_checkbox,\n", " generate_button,\n", " output_area,\n", "])\n", "\n", "display(controls)\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 5. Batch generation (optional)\n", "\n", "Generate several completions of the same prompt at once — useful for comparing outputs at different sampling settings, or just seeing the range of what the model produces.\n" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "def batch_generate(prompt: str, n: int = 5, max_new_tokens: int = 40, temperature: float = 0.5, top_k: int = 40):\n", " for i in range(n):\n", " text = run_generation(prompt, max_new_tokens=max_new_tokens, temperature=temperature, top_k=top_k)\n", " print(f\"[{i + 1}] {text}\\n{'-' * 60}\")\n", "\n", "# Example — feel free to edit the arguments and re-run\n", "batch_generate(\"Deep in the forest, Jack found a shell\", n=3, max_new_tokens=35, temperature=0.5, top_k=40)\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## About this model\n", "\n", "**NoTokenLM-Gen-3.6** is a ~13.9M-parameter, byte-level transformer (RoPE, RMSNorm, SwiGLU, plus a short input convolution, QK-norm, value residual and per-head output gates) trained from scratch on an English + French mix of educational web text, Wikipedia and simple stories. It has no tokenizer — it reads and writes raw UTF-8 bytes directly.\n", "\n", "It writes well-formed English sentences, and it has **no usable world knowledge** — it isn't meant for factual questions, math, code, reasoning, or conversation. See the full [model card on Hugging Face](https://huggingface.co/omurberaisik/NoTokenLM-Gen-3.6) for the complete evaluation results and an honest breakdown of what it can and can't do.\n" ] } ], "metadata": { "colab": { "provenance": [], "name": "NoTokenLM-Gen-3.6-Playground.ipynb" }, "kernelspec": { "display_name": "Python 3", "language": "python", "name": "python3" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 3 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", "version": "3.10.12" }, "accelerator": "GPU" }, "nbformat": 4, "nbformat_minor": 5 }