Run an uncensored AI model on your own computer
A free, offline setup for an abliterated open-weights model with Ollama: install it, pick the size your graphics card can really hold, fix the context-length trap nobody mentions, and keep the port off the internet. Plus what “uncensored” actually means, and where it falls down. Checked against Ollama’s docs, the model’s page and the original research
A model on your own machine has no account attached, no usage meter, no network connection after the download and, if you pick an abliterated one, no list of things it won’t discuss. This is the whole setup with Ollama and an abliterated Qwen 3.5: install the runtime, make sure its server is actually running, choose the largest model your graphics memory can hold, and raise the context window Ollama quietly keeps small. It is honest about the trade. A hosted frontier model is smarter than anything a laptop runs, and removing refusals adds willingness, not capability. The reasons to do it are privacy, offline use, and questions a hosted filter blocks for no good reason.
An open-weights language model, Alibaba’s Qwen 3.5, with its ability to refuse edited out of the weights, running entirely on your own computer through Ollama, a free, MIT-licensed runtime. After the one-time download nothing needs the internet. A hosted frontier model is still smarter; the point is privacy, offline use and no content filter, not quality.
What held up, on 2 October 2026
Checked against the model’s Ollama page and Hugging Face cards, Ollama’s documentation and source, the original refusal-direction paper, Docker’s documentation, and the published scans of exposed Ollama servers.
What “uncensored” actually means here
The model is huihui_ai/qwen3.5-abliterated: Alibaba’s Qwen 3.5, an ordinary open-weights model, put through a process called abliteration.
In 2024 a team of researchers showed that refusal in chat models is carried by a single direction in the model’s internal activations: one direction per model, found across 13 open chat models up to 72 billion parameters. Erase that direction and the model stops refusing; add it and the model refuses even harmless requests. Project it out of the weight matrices and the change is permanent. The open-source community named the result abliteration (ablate plus obliterate) within weeks, and the model cards point to a small proof-of-concept repo, remove-refusals-with-transformers, to show how it works with nothing but the standard Hugging Face library.
So this is surgery on the weights, not a prompt trick. It needs the weights, and it changes the model rather than what you type into it. There is no special prompt to craft: the model answers because the part of it that said no has been removed.
The model’s own page warns that its “safety filtering has been significantly reduced, potentially generating sensitive, controversial, or inappropriate content,” recommends it for “research, testing, or controlled environments,” and says users are “solely responsible for any consequences.” Read that as the authors telling you the truth rather than covering themselves.
Install Ollama
Ollama is the runtime: it downloads models, loads them into memory and gives you a prompt to type at. It is the easiest part of this.
Linux
One command, below. It installs the binary, creates an ollama service user and a systemd service that starts on boot, and reports the API at 127.0.0.1:11434. It needs zstd to unpack; install that first if the script asks.
macOS
Download the .dmg from ollama.com/download and drag Ollama to Applications, or use the same one-line script. Needs macOS 14 Sonoma or newer. Apple Silicon runs models on the GPU through Metal; Intel Macs run on the CPU only, which matters a lot below.
Windows
Download OllamaSetup.exe from ollama.com/download, or run the PowerShell line below. Needs Windows 10 22H2 or newer. It doesn’t ask for Administrator and installs into your home folder, so you need at least 4 GB free there before any models. NVIDIA cards want driver 551.61 or newer.
Install
curl | sh runs a script from the internet in your shell. It is the method Ollama publishes; to read it first, open ollama.com/install.sh in a browser.
Make sure the server is actually running
This is the step that gets skipped, and it is the one behind most confused messages. Ollama is two things: a background server that holds models in memory, and a command-line client that talks to it. On macOS and Windows the app runs the server from the menu bar or system tray, and a command will start the app if it isn’t running. On Linux the install script sets up a systemd service. Install by hand, stop the service or work in a container, and nothing is listening: every command then fails with the same error, and the fix is in the error.
Start it, and check it answers
Pick a size, and get this one right
This is the decision the whole thing hinges on. Open the model’s page on Ollama and look at its tags. The number is parameters in billions, and the download size next to it is roughly what has to fit in memory.
The rule: take your graphics card’s memory, subtract a gigabyte or two for the system and the conversation, and pick the largest tag under what’s left. On a 16 GB card the 9b (6.6 GB) is comfortable and the 27b (17 GB) is not; its weights alone are bigger than the card. To find your number: on Windows, Task Manager → Performance → GPU → Dedicated GPU memory; on Linux, nvidia-smi. Macs work differently, below.
The sizes
Download sizes from the model’s Ollama page; the memory column is a rule of thumb, the download plus room for context. The 35b and 122B are mixture-of-experts models (3B and 10B parameters active per token), which helps speed, not memory: every parameter still has to fit.
When it doesn’t fit, nothing dramatic happens. That’s the problem
Ollama doesn’t refuse a model that is too big for the graphics card. It puts what it can on the GPU and the rest in system RAM, where the CPU does the work. The model still answers, just several times slower, and most people assume that is simply how local AI feels and give up. You can see exactly what happened: with the model loaded, run ollama ps.
Where did it load?
Illustrative output. SIZE is the memory in use, weights plus context, so it runs above the download size. CONTEXT is the window you actually got.
Reading the PROCESSOR column
It is the whole story. If it reads anything but 100% GPU and you care about speed, drop one size and pull again.
No graphics card, or a Mac
No dedicated graphics card still works: Ollama falls back to the CPU and system RAM. A 2B or 4B model on a modern CPU is slow but usable for short answers. A 9b on the CPU alone is slow enough to test anyone’s patience.
Apple Silicon is the exception worth knowing. The CPU and GPU share one pool of memory, so a Mac can hold models that would need an expensive graphics card on a PC. The GPU can’t use all of it: Apple’s own figures are 21 GB of a 32 GB machine and 48 GB of a 64 GB one, roughly two-thirds to three-quarters. So a 16 GB Mac handles the 9b comfortably, and a 32 GB Mac can fit the 27b.
Context length
The model’s page says 256K context. You won’t get that by default. Ollama sizes the default to your graphics memory, and under 24 GiB it gives you 4,096 tokens: roughly three thousand words of memory in total, the model’s own replies included. That is why a long conversation starts forgetting its beginning, and it isn’t the model being dumb. From 24 to 48 GiB you get 32k, and only 48 GiB and up gets the full 256k. On a Mac the figure that counts is what the GPU may use, so even a 32 GB Mac probably lands in the 4k tier; the server log’s “vram-based default context” line tells you which you got.
Raise it
A bigger window costs memory on top of the weights, so if you push the context up, drop the model size down. Ollama’s docs ask for at least 64,000 tokens for agents, coding tools and web search, which realistically means a 24 GB card and a small model.
Pull it, run it
Swap 4B for the tag you picked, and always give a tag: with none, Ollama pulls the default, which for this model is the 17 GB 27b. You’ll see the layers download, a sha256 check, and success.
Then turn your wifi off and keep going. Nothing here needs the internet after the download, and it is worth doing once just to watch it work. Ollama also offers optional cloud models and web search; to rule those out entirely, start the server with OLLAMA_NO_CLOUD=1.
Where the models land
Worth knowing when you need the space back.
The Docker version
Same thing, isolated and easy to delete. This is the NVIDIA form; the setup for each kind of hardware follows.
GPU setup by vendor
The command you’ll find everywhere uses -p 11434:11434, and Docker publishes that on every network interface, straight past ufw. Inside the official image Ollama listens on all addresses, so the host address in -p is the only thing keeping the API off your network. Bind it to 127.0.0.1 so only your machine can reach it, and pick an unusual outside port (43219 here; choose your own) instead of 11434, the port scanners look for. The API has no authentication, and exposed servers number in the tens of thousands. The container still speaks 11434 inside; only the outside number changes.
Troubleshooting
Willingness is not capability
A 4B model is not a frontier model, and the gap is enormous. It will write confident code that doesn’t compile, invent library functions, and lose track of what you asked three messages ago. Removing refusals adds no capability; the research measures a small cost, mostly on truthfulness benchmarks.
The telling test is the one every “uncensored AI writes malware” clip runs. Asked for malware to steal someone’s Bitcoin, the 4B agreed at once, cheerfully, then produced a program that sends a Bitcoin transaction with a private key you have to type in yourself. That is a wallet, badly written. The willingness was real and the capability was not, and that is the shape of almost every one of those clips.
What it is genuinely good for
Whatever you generate locally is subject to the same laws as anything else you write, and “a model made it” has never been a defence. And if you expose the API port to the internet, you have handed a stranger free use of your graphics card. Keep it on localhost.
Install Ollama. Run ollama serve if nothing is listening. Check your graphics memory, pick the tag one size under it, pull it with the tag, run it, and check that ollama ps reads 100% GPU. Then turn the wifi off and ask it something.