walid@portfolio:~/lab/local-uncensored-ai$
cd../lab
01ideaOct 2026

Run an uncensored AI model on your own computer

A free, offline setup for an abliterated open-weights model with Ollama: install it, pick the size your graphics card can really hold, fix the context-length trap nobody mentions, and keep the port off the internet. Plus what “uncensored” actually means, and where it falls down. Checked against Ollama’s docs, the model’s page and the original research

A model on your own machine has no account attached, no usage meter, no network connection after the download and, if you pick an abliterated one, no list of things it won’t discuss. This is the whole setup with Ollama and an abliterated Qwen 3.5: install the runtime, make sure its server is actually running, choose the largest model your graphics memory can hold, and raise the context window Ollama quietly keeps small. It is honest about the trade. A hosted frontier model is smarter than anything a laptop runs, and removing refusals adds willingness, not capability. The reasons to do it are privacy, offline use, and questions a hosted filter blocks for no good reason.

OllamaLocal AIOpen weightsQwen 3.5AbliterationPrivacyOllama ↗Ollama download ↗Ollama docs ↗The model on Ollama ↗The refusal-direction paper ↗remove-refusals-with-transformers ↗
i
What this is

An open-weights language model, Alibaba’s Qwen 3.5, with its ability to refuse edited out of the weights, running entirely on your own computer through Ollama, a free, MIT-licensed runtime. After the one-time download nothing needs the internet. A hosted frontier model is still smarter; the point is privacy, offline use and no content filter, not quality.

Checked, not copied

What held up, on 2 October 2026

Checked against the model’s Ollama page and Hugging Face cards, Ollama’s documentation and source, the original refusal-direction paper, Docker’s documentation, and the published scans of exposed Ollama servers.

ClaimWhat the source says
The model and its reachReal. huihui_ai/qwen3.5-abliterated is live on Ollama with about 657,000 downloads (657.1K on 2 October), last updated 5 April 2026. Every tag lists a 256K context window and text-and-image input. Qwen 3.5 itself is Apache-2.0, released February–March 2026, and no longer Qwen’s newest family.
The size tableExact: 0.8B 1.0 GB, 2B 1.9 GB, 4B 3.3 GB, 9b 6.6 GB, 27b 17 GB, 35b 24 GB, 122B 81 GB. One gotcha the guide misses: the untagged default (“latest”) is the 27b, so a bare ollama run with no tag downloads 17 GB.
What the 4B isNot quite what the page implies. Its Hugging Face card says the 4B was uncensored by fine-tuning (trained with TRL); the other sizes say abliteration. Same goal, different method.
The researchReal: Arditi and colleagues, “Refusal in Language Models Is Mediated by a Single Direction” (NeurIPS 2024), found one refusal direction in each of 13 open chat models up to 72B, and showed it can be projected out of the weights. The paper calls that a “white-box jailbreak”, so “not a prompt trick” is right and “not jailbreaking” is not.
remove-refusals-with-transformersReal, and the model’s own readme points to it to explain the method. The repo is a self-described “crude, proof-of-concept” that removes the direction while the model runs; it doesn’t save edited weights itself.
“Abliteration degrades a model slightly”Supported. In the paper, MMLU, ARC and GSM8K stayed within noise for most models while TruthfulQA dropped every time; an independent write-up measured a small drop across benchmarks. Nobody has published numbers for this particular model.
Ollama’s requirementsMatch the docs: MIT licence; macOS 14 Sonoma or newer, Apple Silicon on the GPU and Intel on the CPU only; Windows 10 22H2 or newer, no Administrator, at least 4 GB for the install, NVIDIA driver 551.61 or newer.
The 4,096-token defaultCorrect, and now tiered by graphics memory: under 24 GiB Ollama defaults to 4k tokens, 24–48 GiB to 32k, 48 GiB and up to 256k. The docs ask for at least 64,000 tokens for agents and coding tools.
“Ten to thirty times slower”No source quantifies it. Ollama’s docs only say responses “may be slower” once a model spills into system RAM; the hit depends on how much spills and on your memory bandwidth.
Exposed serversReal and growing. Wiz found over 1,000 exposed in 2024 and Cisco 1,139 in September 2025; SentinelLABS with Censys logged 175,108 distinct Ollama hosts across 293 days of scanning (January 2026). Ollama’s local API has no authentication.
The ollama ps sampleOut of date. The current CLI adds a CONTEXT column, and the 4B’s ID is 9acbd90c2645; the sample showed the hash of its weights file instead. Corrected below.
Before you install

What “uncensored” actually means here

The model is huihui_ai/qwen3.5-abliterated: Alibaba’s Qwen 3.5, an ordinary open-weights model, put through a process called abliteration.

In 2024 a team of researchers showed that refusal in chat models is carried by a single direction in the model’s internal activations: one direction per model, found across 13 open chat models up to 72 billion parameters. Erase that direction and the model stops refusing; add it and the model refuses even harmless requests. Project it out of the weight matrices and the change is permanent. The open-source community named the result abliteration (ablate plus obliterate) within weeks, and the model cards point to a small proof-of-concept repo, remove-refusals-with-transformers, to show how it works with nothing but the standard Hugging Face library.

So this is surgery on the weights, not a prompt trick. It needs the weights, and it changes the model rather than what you type into it. There is no special prompt to craft: the model answers because the part of it that said no has been removed.

!
The publisher says so, plainly

The model’s own page warns that its “safety filtering has been significantly reduced, potentially generating sensitive, controversial, or inappropriate content,” recommends it for “research, testing, or controlled environments,” and says users are “solely responsible for any consequences.” Read that as the authors telling you the truth rather than covering themselves.

Step 1

Install Ollama

Ollama is the runtime: it downloads models, loads them into memory and gives you a prompt to type at. It is the easiest part of this.

1

Linux

One command, below. It installs the binary, creates an ollama service user and a systemd service that starts on boot, and reports the API at 127.0.0.1:11434. It needs zstd to unpack; install that first if the script asks.

2

macOS

Download the .dmg from ollama.com/download and drag Ollama to Applications, or use the same one-line script. Needs macOS 14 Sonoma or newer. Apple Silicon runs models on the GPU through Metal; Intel Macs run on the CPU only, which matters a lot below.

3

Windows

Download OllamaSetup.exe from ollama.com/download, or run the PowerShell line below. Needs Windows 10 22H2 or newer. It doesn’t ask for Administrator and installs into your home folder, so you need at least 4 GB free there before any models. NVIDIA cards want driver 551.61 or newer.

Install

curl | sh runs a script from the internet in your shell. It is the method Ollama publishes; to read it first, open ollama.com/install.sh in a browser.

install-ollama5 lines
# Linux or macOS
curl -fsSL https://ollama.com/install.sh | sh

# Windows (PowerShell)
irm https://ollama.com/install.ps1 | iex
Step 1.5

Make sure the server is actually running

This is the step that gets skipped, and it is the one behind most confused messages. Ollama is two things: a background server that holds models in memory, and a command-line client that talks to it. On macOS and Windows the app runs the server from the menu bar or system tray, and a command will start the app if it isn’t running. On Linux the install script sets up a systemd service. Install by hand, stop the service or work in a container, and nothing is listening: every command then fails with the same error, and the fix is in the error.

Start it, and check it answers

start-the-server11 lines
# on Linux, when nothing is listening, every command stops with:
#   Error: could not connect to ollama server, run 'ollama serve' to start it

# terminal 1: start the server and leave it open (its log prints here)
ollama serve

# terminal 2: check that it answers
ollama -v

# or, on Linux with systemd, start the service instead
sudo systemctl start ollama
Step 2

Pick a size, and get this one right

This is the decision the whole thing hinges on. Open the model’s page on Ollama and look at its tags. The number is parameters in billions, and the download size next to it is roughly what has to fit in memory.

The rule: take your graphics card’s memory, subtract a gigabyte or two for the system and the conversation, and pick the largest tag under what’s left. On a 16 GB card the 9b (6.6 GB) is comfortable and the 27b (17 GB) is not; its weights alone are bigger than the card. To find your number: on Windows, Task Manager → Performance → GPU → Dedicated GPU memory; on Linux, nvidia-smi. Macs work differently, below.

The sizes

Download sizes from the model’s Ollama page; the memory column is a rule of thumb, the download plus room for context. The 35b and 122B are mixture-of-experts models (3B and 10B parameters active per token), which helps speed, not memory: every parameter still has to fit.

TagDownloadGraphics memory you need
0.8BDownload1.0 GBGraphics memory you need2 GB
2BDownload1.9 GBGraphics memory you need4 GB
4BDownload3.3 GBGraphics memory you need6 GB
9bDownload6.6 GBGraphics memory you need8 GB
27bDownload17 GBGraphics memory you need20 GB
35bDownload24 GBGraphics memory you need28 GB
122BDownload81 GBGraphics memory you needNot your laptop

When it doesn’t fit, nothing dramatic happens. That’s the problem

Ollama doesn’t refuse a model that is too big for the graphics card. It puts what it can on the GPU and the rest in system RAM, where the CPU does the work. The model still answers, just several times slower, and most people assume that is simply how local AI feels and give up. You can see exactly what happened: with the model loaded, run ollama ps.

Where did it load?

Illustrative output. SIZE is the memory in use, weights plus context, so it runs above the download size. CONTEXT is the window you actually got.

ollama-ps4 lines
ollama ps

NAME                                ID              SIZE      PROCESSOR    CONTEXT    UNTIL
huihui_ai/qwen3.5-abliterated:4B    9acbd90c2645    4.1 GB    100% GPU     4096       4 minutes from now

Reading the PROCESSOR column

It is the whole story. If it reads anything but 100% GPU and you care about speed, drop one size and pull again.

PROCESSOR readsWhat it means
100% GPUEverything fit. You are getting the speed the hardware can give.
100% CPUThe graphics card was never used; the model is running in system RAM.
48%/52% CPU/GPUSplit across both, which is why it crawls. The percentages are shares of the model’s memory.

No graphics card, or a Mac

No dedicated graphics card still works: Ollama falls back to the CPU and system RAM. A 2B or 4B model on a modern CPU is slow but usable for short answers. A 9b on the CPU alone is slow enough to test anyone’s patience.

Apple Silicon is the exception worth knowing. The CPU and GPU share one pool of memory, so a Mac can hold models that would need an expensive graphics card on a PC. The GPU can’t use all of it: Apple’s own figures are 21 GB of a 32 GB machine and 48 GB of a 64 GB one, roughly two-thirds to three-quarters. So a 16 GB Mac handles the 9b comfortably, and a 32 GB Mac can fit the 27b.

The trap nobody mentions

Context length

The model’s page says 256K context. You won’t get that by default. Ollama sizes the default to your graphics memory, and under 24 GiB it gives you 4,096 tokens: roughly three thousand words of memory in total, the model’s own replies included. That is why a long conversation starts forgetting its beginning, and it isn’t the model being dumb. From 24 to 48 GiB you get 32k, and only 48 GiB and up gets the full 256k. On a Mac the figure that counts is what the GPU may use, so even a 32 GB Mac probably lands in the 4k tier; the server log’s “vram-based default context” line tells you which you got.

Raise it

raise-context5 lines
# raise the default for everything the server loads
OLLAMA_CONTEXT_LENGTH=32768 ollama serve

# or for one conversation, typed inside the chat
/set parameter num_ctx 32768

A bigger window costs memory on top of the weights, so if you push the context up, drop the model size down. Ollama’s docs ask for at least 64,000 tokens for agents, coding tools and web search, which realistically means a 24 GB card and a small model.

Steps 3 and 4

Pull it, run it

Swap 4B for the tag you picked, and always give a tag: with none, Ollama pulls the default, which for this model is the 17 GB 27b. You’ll see the layers download, a sha256 check, and success.

pull-and-run5 lines
ollama pull huihui_ai/qwen3.5-abliterated:4B
ollama run huihui_ai/qwen3.5-abliterated:4B

# loaded when you see:  >>> Send a message (/? for help)
# /bye exits · /? lists the commands

Then turn your wifi off and keep going. Nothing here needs the internet after the download, and it is worth doing once just to watch it work. Ollama also offers optional cloud models and web search; to rule those out entirely, start the server with OLLAMA_NO_CLOUD=1.

Where the models land

Worth knowing when you need the space back.

SystemLocation
Linux/usr/share/ollama/.ollama/models
macOS~/.ollama/models
WindowsC:\Users\<you>\.ollama\models
Somewhere elseSet OLLAMA_MODELS to a folder on another drive. On Linux the ollama service user has to own it: sudo chown -R ollama:ollama <folder>.
Optional

The Docker version

Same thing, isolated and easy to delete. This is the NVIDIA form; the setup for each kind of hardware follows.

docker-run6 lines
docker run -d --gpus=all \
  -v ollama:/root/.ollama \
  -p 127.0.0.1:43219:11434 \
  --name ollama ollama/ollama

docker exec -it ollama ollama run huihui_ai/qwen3.5-abliterated:4B

GPU setup by vendor

docker-gpu11 lines
# NVIDIA: install the NVIDIA Container Toolkit first, then
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

# CPU only: the same run command without --gpus=all

# AMD: the rocm image, with the GPU devices passed through
docker run -d --device /dev/kfd --device /dev/dri \
  -v ollama:/root/.ollama \
  -p 127.0.0.1:43219:11434 \
  --name ollama ollama/ollama:rocm
!
Two things about that port mapping

The command you’ll find everywhere uses -p 11434:11434, and Docker publishes that on every network interface, straight past ufw. Inside the official image Ollama listens on all addresses, so the host address in -p is the only thing keeping the API off your network. Bind it to 127.0.0.1 so only your machine can reach it, and pick an unusual outside port (43219 here; choose your own) instead of 11434, the port scanners look for. The API has no authentication, and exposed servers number in the tens of thousands. The container still speaks 11434 inside; only the outside number changes.

When it breaks

Troubleshooting

SymptomFix
It answers slowlyRun ollama ps and read PROCESSOR. Anything but 100% GPU means the model spilled into system RAM: drop a size. This covers most complaints.
Nothing connectsThe server isn’t running: ollama serve, or sudo systemctl start ollama.
The download was 17 GBYou ran it without a tag, which pulls the 27b. Add :4B, :9b or whichever tag you picked.
The install script wants zstdCurrent Linux packages are .tar.zst. Install zstd (sudo apt-get install zstd on Debian or Ubuntu) and run it again.
The GPU vanished after sleepA known NVIDIA driver issue on Linux after suspend and resume: sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm brings it back.
You want the logsLinux: journalctl -u ollama --no-pager --follow --pager-end. macOS: cat ~/.ollama/logs/server.log. Windows: %LOCALAPPDATA%\Ollama\server.log. Docker: docker logs ollama. Started by hand with ollama serve: that terminal.
Where this falls down

Willingness is not capability

A 4B model is not a frontier model, and the gap is enormous. It will write confident code that doesn’t compile, invent library functions, and lose track of what you asked three messages ago. Removing refusals adds no capability; the research measures a small cost, mostly on truthfulness benchmarks.

The telling test is the one every “uncensored AI writes malware” clip runs. Asked for malware to steal someone’s Bitcoin, the 4B agreed at once, cheerfully, then produced a program that sends a Bitcoin transaction with a private key you have to type in yourself. That is a wallet, badly written. The willingness was real and the capability was not, and that is the shape of almost every one of those clips.

What it is genuinely good for

✓Security research, where ordinary questions trip a hosted filter for no reason✓Fiction with content a hosted model won’t touch✓Working on a plane, or anywhere without a connection✓Anything you would rather not upload to a company
!
Two limits that aren’t technical

Whatever you generate locally is subject to the same laws as anything else you write, and “a model made it” has never been a defence. And if you expose the API port to the internet, you have handed a stranger free use of your graphics card. Keep it on localhost.

i
The short version

Install Ollama. Run ollama serve if nothing is listening. Check your graphics memory, pick the tag one size under it, pull it with the tag, run it, and check that ollama ps reads 100% GPU. Then turn the wifi off and ask it something.