20 August 2026

How to Turn Kaggle Into a Free AI Server (And Accidentally Invite the Entire Internet)

SARCASM WARNING! (Also: code warning. There’s actual working stuff buried in here, so don’t skim too hard.)

Well hello there, my GPU-starved friends! Tired of watching your laptop wheeze like an asthmatic hamster every time you try to run a real AI model? Good news. Kaggle will hand you two 16 GB GPUs, about 30 hours a week, for the low, low price of absolutely nothing.

Here’s the math. A 27-billion-parameter language model needs about 17 GB of GPU memory just to wake up. Your laptop has 8 GB of RAM, a sticker from a 2019 conference, and dreams. Kaggle has the hardware. You have the audacity. Let’s make a deal.

Here’s the whole trick: take a Kaggle notebook (built for data science homework), turn it into a private AI server, load an “uncensored” Qwen model onto it, punch a hole to the internet with a Cloudflare tunnel, and point Claude Code at it from your own machine. Fifteen minutes. Zero dollars. What could possibly go wrong?

Spoiler: one thing. One very large, very open, very front-door-shaped thing. I went through every command so you don’t have to, fixed the ones that quietly break, and found the problem nobody seems to mention. It’s the most important part of this post, so naturally I’ve put it near the end, where nobody reads.


🎁 Kaggle: A GPU Rental Shop That Forgot to Charge You

Kaggle exists so people can enter machine-learning competitions and argue about leaderboards. To keep that fair, it gives every verified account free notebook time on real GPUs. Real ones. Not the “integrated graphics” kind your IT department calls a workstation.

The option you want is GPU T4 x2: two NVIDIA T4 cards with about 15 GB of usable memory each.

Why two? Because the model, a 4-bit build of Qwen3.8-27B, weighs roughly 17 GB, and one T4 can’t hold it. Kaggle’s other option, a single P100, can’t either once you leave room for the actual conversation. Two T4s can, and Ollama splits the model across them without being asked. Teamwork. Something your SOC and dev team could learn from.

Before anything works, flip two switches:

  • Verify your phone number in Kaggle settings. No phone, no GPU, no internet. Kaggle has trust issues. Respect.
  • In the notebook’s Session options, set Accelerator to GPU T4 x2 and Internet to On.

Now the fine print, because there’s always fine print. Sessions die after about 12 hours or when idle. The weekly quota is about 30 GPU hours. And every new session starts from absolute zero: a fresh 17 GB download and a brand-new public URL. It’s Groundhog Day, but with CUDA. Remember this. It comes back to bite later.


🪤 Four Commands, Three Traps

The heart of all this is Ollama, a tool that downloads open models and serves them over a local API. Getting it running on Kaggle takes four steps. Three of them are booby-trapped, and the obvious way of doing it strolls right over them like a tourist in a minefield. Here’s the damage report:

StepThe obvious wayWhat actually happensSanity level
Install OllamaRuns the official install scriptFails: Kaggle has no zstdMildly shaken
Start OllamaTrusts the defaults4,096-token goldfish memorySuspicious
Get the modelollama runCell waits forever for someone to typePhilosophically broken
Open the tunnelReads 30 lines of log, then stopsTunnel can freeze hours laterWho am I anymore?
Secure itAPI key: ollamaProtects absolutely nothingNumb but enlightened

Trap one: the installer needs a tool Kaggle doesn’t have. Ollama ships as a .tar.zst archive, and Kaggle’s image has no zstd to unpack it. That’s IKEA delivering your wardrobe without the Allen key. Install it first:

!sudo apt-get install -y zstd
!curl -fsSL https://ollama.com/install.sh | sh

The installer will then complain that systemd isn’t running. Ignore it. Notebooks have no service manager, so you start the server yourself, like an animal.

Trap two: the default memory is goldfish-sized. Ollama gives every model a 4,096-token context window unless told otherwise. Sounds like plenty, until you learn that Claude Code’s opening instructions alone are longer. And here’s the kicker: Ollama doesn’t throw an error when a prompt overflows. It quietly chops off the beginning and carries on, smiling. Your agent forgets what it was doing and you never find out why. It’s walking into a room and forgetting why you came in, except it happens on every single message. So start the server with a bigger window:

import os, subprocess, time

env = os.environ.copy()
env["OLLAMA_CONTEXT_LENGTH"] = "32768"
env["OLLAMA_KEEP_ALIVE"] = "-1"
subprocess.Popen(["ollama", "serve"], env=env,
                 stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
time.sleep(5)

And yes, Popen matters. Run !ollama serve directly and the cell blocks forever, just sitting there. Staring. Judging.

Trap three: run waits for a person who isn’t there. The obvious way to download the model is ollama run. That command downloads, then opens an interactive chat and waits for you to type. A notebook cell has no keyboard. So it waits. And waits. Like a Tamagotchi nobody’s feeding. Use pull, which downloads and actually leaves:

!ollama pull hf.co/JonathanColetti/Qwen3.8-27B-Uncensored-GGUF:Q4_K_M

That odd-looking name is Ollama pulling a GGUF file straight from Hugging Face. Q4_K_M is the 4-bit version, the usual trade-off between size and quality. Think of it as the model’s travel-size shampoo.

😈 What “Uncensored” Actually Means (Calm Down)

Before anyone gets excited: “uncensored” does not mean smarter, edgier, or secretly aware of where the bodies are buried. The model was made with a technique called abliteration, using a tool named Heretic (yes, really). It finds the direction inside the model that produces refusals and removes it, with no retraining. On the author’s 100-prompt test of harmful requests, refusals fell from 98 to 12.

That’s it. That’s the whole upgrade. It is not smarter, and removing pieces of a model usually costs it a little accuracy. It’s like taking the brakes off a car and calling it a sports car. If you want a coding assistant and don’t care about refusals, the stock Qwen model is the better pick.


🚇 A Tunnel Out of a Building With No Doors

Congratulations, you now have a 27B model running on 127.0.0.1:11434, inside a container, inside a Google data centre. Nothing outside can reach it. Kaggle notebooks accept no incoming connections. Your shiny new AI server is basically a genius locked in a basement.

Enter the Cloudflare quick tunnel, which solves this by going the other way. The cloudflared program inside the notebook dials out to Cloudflare and holds the line open. Cloudflare hands you a public address like https://some-random-words.trycloudflare.com, and requests to that address travel down the line to Ollama. No account, no domain, no firewall rules, no change request, no CAB meeting. Your GRC team would faint.

Your requests reach the model on Kaggle through a Cloudflare tunnel Kaggle notebook, GPU T4 x2 Your computer Claude Code OpenAI-style apps Cloudflare public HTTPS URL trycloudflare.com cloudflared outbound only rewrites Host Token proxy port 8080 optional Ollama port 11434 Qwen3.8 27B Q4_K_M Requests run left to right and replies stream back. The tunnel dials out from Kaggle, so no inbound port is opened.
import re, subprocess, threading

cloudflared = subprocess.Popen(
    ["cloudflared", "tunnel", "--url", "http://127.0.0.1:11434",
     "--http-host-header", "localhost:11434"],
    stdout=subprocess.PIPE, stderr=subprocess.STDOUT, text=True)

for line in cloudflared.stdout:
    m = re.search(r"https://[a-z0-9-]+\.trycloudflare\.com", line)
    if m:
        print("Your URL:", m.group(0))
        break
threading.Thread(target=lambda: [None for _ in cloudflared.stdout], daemon=True).start()

Two details here will save you an afternoon of swearing at your screen:

  • --http-host-header localhost:11434 disguises outside requests as local ones. Ollama refuses anything that doesn’t claim to come from localhost, so without this flag every call comes back 403 Forbidden. Ollama is the bouncer; this flag is the fake ID.
  • The last line keeps reading the tunnel’s log in the background. The obvious version reads 30 lines and then stops listening, like a manager in a risk review. cloudflared keeps talking anyway, the buffer fills up, and the tunnel can freeze hours later with no error. Silent failures: the best kind.

🤖 Plugging In Claude Code

Since version 0.14, Ollama speaks Anthropic’s API as well as OpenAI’s. Bilingual. Overachiever. That means Claude Code can talk to it directly:

export ANTHROPIC_BASE_URL="https://some-random-words.trycloudflare.com"
export ANTHROPIC_AUTH_TOKEN="ollama"
export ANTHROPIC_API_KEY=""
claude --model hf.co/JonathanColetti/Qwen3.8-27B-Uncensored-GGUF:Q4_K_M

Now notice what’s missing: /v1. OpenAI-style apps such as Open WebUI or Cursor want https://…trycloudflare.com/v1. Claude Code adds its own /v1/messages, so if you give it /v1 too, you get /v1/v1/messages and a 404. One slash. ONE. It’s the easiest mistake in the whole setup to make and the hardest to spot, which, coincidentally, describes most of IT.

You’ll also see people rename the model to claude-sonnet-4-5 with ollama cp, so Claude Code finds a familiar name. It works, but passing the real name with --model does the same job without the cosplay.

And let’s manage expectations, shall we? A 4-bit 27B model on two T4s is not a hosted Claude model in a fake moustache. It types slowly and gets lost in long, multi-step coding tasks. Great for experiments. Terrible as your daily driver. Like a free gym membership: technically it works, but we both know how this ends.


🚪 The Front Door Is Wide Open (The Part Everyone Skips)

And now, the moment you’ve been scrolling for.

The API key in this setup is the word ollama. It looks like a password. It is not a password. It is a word. Clients refuse to start without some key, so you type a placeholder, and Ollama never checks it. Ollama has no authentication. None. Zero. Zilch. It’s a nightclub with no bouncer, no door, and a neon sign that says FREE GPUs.

So the moment that tunnel comes up, your model sits on the public internet behind a URL and absolutely nothing else. Anyone who finds that URL can:

  • burn your weekly GPU hours on their own requests (thanks for the free compute, buddy);
  • call /api/delete and wipe your model, or /api/pull and fill your disk;
  • send whatever they like through an uncensored model, with your account attached. Enjoy explaining that one.

“But the URL is random!” I hear you cry. Sure. Until you paste it into a screenshot, a Discord message, a shared notebook, or a log file. A random address hides the door. It doesn’t lock it. Security through obscurity: the corporate classic, now available for hobbyists.

🔐 The Fix: A Tiny Proxy With a Password

Put a small gatekeeper between the tunnel and Ollama. It checks every request for a secret token and slams the door on anything without one. Claude Code already sends ANTHROPIC_AUTH_TOKEN as a bearer token, and OpenAI-style apps send their API key the same way, so your clients need zero changes beyond swapping ollama for your secret. Look at that: actual security, no 200-page policy document required.

%%writefile proxy.py
import os, aiohttp
from aiohttp import web

SECRET = os.environ["PROXY_TOKEN"]
UPSTREAM = "http://127.0.0.1:11434"

async def handle(request):
    if (request.headers.get("Authorization") != f"Bearer {SECRET}"
            and request.headers.get("x-api-key") != SECRET):
        return web.Response(status=401, text="unauthorized")
    headers = {k: v for k, v in request.headers.items()
               if k.lower() not in ("host", "authorization", "x-api-key", "content-length")}
    async with aiohttp.ClientSession(timeout=aiohttp.ClientTimeout(total=None)) as s:
        async with s.request(request.method, UPSTREAM + request.rel_url.path_qs,
                             headers=headers, data=await request.read()) as up:
            resp = web.StreamResponse(status=up.status, headers={
                "Content-Type": up.headers.get("Content-Type", "application/json")})
            await resp.prepare(request)
            async for chunk in up.content.iter_any():
                await resp.write(chunk)
            await resp.write_eof()
            return resp

app = web.Application(client_max_size=50 * 1024**2)
app.router.add_route("*", "/{tail:.*}", handle)
web.run_app(app, host="127.0.0.1", port=8080, print=None)

Start it with a random secret, then point the tunnel at port 8080 instead of 11434:

import secrets
TOKEN = secrets.token_urlsafe(32)
subprocess.Popen(["python", "proxy.py"], env={**os.environ, "PROXY_TOKEN": TOKEN})
print("Your API key:", TOKEN)

Replies still stream token by token, and the secret is stripped out before anything reaches Ollama. If you own a domain, a named Cloudflare tunnel with Cloudflare Access does the same check at Cloudflare’s edge. Either way, when you’re done, close the door behind you: cloudflared.terminate().

One more thing, before you build your entire personality around this. Kaggle gives out free GPUs for data science and ML work. I couldn’t find a rule that bans tunnels outright, but running a public model server is clearly not what the free tier is for. Read the terms before you make it part of your routine. “I didn’t read the terms” has never once worked as a defence.


🏁 Final Rant: Free GPUs, Borrowed Time

So, is it worth it? For the right job, absolutely.

If you want to try a 27B model before buying hardware, compare an abliterated model with its stock version, or see how far an open model gets inside Claude Code, this is the cheapest lab on the planet. Fifteen minutes and zero dollars gets you a model most laptops can’t even load.

But know what you’re holding. It’s a server that vanishes every twelve hours and changes its address every time it comes back, like a witness in protection. Its model is a step down from the hosted ones. And its front door is open until you lock it. If you need something always on, rent a GPU by the hour from RunPod or Vast.ai and get a fixed address. If you need privacy, run Ollama on your own machine.

But all sarcasm aside, the real lesson here isn’t the tunnel or the model. It’s that putting something on the internet now takes one command, and securing it still takes thought. The command is the easy part. Bring the thought yourself.

Because at the end of the day, free GPUs don’t fix stupid either.

Happy (authenticated) tunneling!