{"id":1000024,"date":"2026-08-20T12:48:07","date_gmt":"2026-08-20T11:48:07","guid":{"rendered":"https:\/\/siyaz.tech\/?p=1000024"},"modified":"2026-10-11T13:37:29","modified_gmt":"2026-10-11T12:37:29","slug":"kaggle-free-ai-server-ollama-cloudflare-tunnel","status":"publish","type":"post","link":"https:\/\/siyaz.tech\/index.php\/2026\/08\/20\/kaggle-free-ai-server-ollama-cloudflare-tunnel\/","title":{"rendered":"How to Turn Kaggle Into a Free AI Server (And Accidentally Invite the Entire Internet)"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">SARCASM WARNING! (Also: code warning. There&#8217;s actual working stuff buried in here, so don&#8217;t skim too hard.)<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Well hello there, my GPU-starved friends! Tired of watching your laptop wheeze like an asthmatic hamster every time you try to run a real AI model? Good news. Kaggle will hand you two 16 GB GPUs, about 30 hours a week, for the low, low price of absolutely nothing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Here&#8217;s the math. A 27-billion-parameter language model needs about 17 GB of GPU memory just to wake up. Your laptop has 8 GB of RAM, a sticker from a 2019 conference, and dreams. Kaggle has the hardware. You have the audacity. Let&#8217;s make a deal.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Here&#8217;s the whole trick: take a Kaggle notebook (built for data science homework), turn it into a private AI server, load an &#8220;uncensored&#8221; Qwen model onto it, punch a hole to the internet with a Cloudflare tunnel, and point Claude Code at it from your own machine. Fifteen minutes. Zero dollars. What could possibly go wrong?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Spoiler: one thing. One very large, very open, very front-door-shaped thing. I went through every command so you don&#8217;t have to, fixed the ones that quietly break, and found the problem nobody seems to mention. It&#8217;s the most important part of this post, so naturally I&#8217;ve put it near the end, where nobody reads.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83c\udf81 Kaggle: A GPU Rental Shop That Forgot to Charge You<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Kaggle exists so people can enter machine-learning competitions and argue about leaderboards. To keep that fair, it gives every verified account free notebook time on real GPUs. Real ones. Not the &#8220;integrated graphics&#8221; kind your IT department calls a workstation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The option you want is <strong>GPU T4 x2<\/strong>: two NVIDIA T4 cards with about 15 GB of usable memory each.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Why two? Because the model, a 4-bit build of Qwen3.8-27B, weighs roughly 17 GB, and one T4 can&#8217;t hold it. Kaggle&#8217;s other option, a single P100, can&#8217;t either once you leave room for the actual conversation. Two T4s can, and Ollama splits the model across them without being asked. Teamwork. Something your SOC and dev team could learn from.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Before anything works, flip two switches:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Verify your phone number<\/strong> in Kaggle settings. No phone, no GPU, no internet. Kaggle has trust issues. Respect.<\/li>\n\n\n<li>In the notebook&#8217;s <strong>Session options<\/strong>, set Accelerator to <strong>GPU T4 x2<\/strong> and Internet to <strong>On<\/strong>.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Now the fine print, because there&#8217;s always fine print. Sessions die after about 12 hours or when idle. The weekly quota is about 30 GPU hours. And every new session starts from absolute zero: a fresh 17 GB download and a brand-new public URL. It&#8217;s Groundhog Day, but with CUDA. Remember this. It comes back to bite later.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83e\udea4 Four Commands, Three Traps<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The heart of all this is Ollama, a tool that downloads open models and serves them over a local API. Getting it running on Kaggle takes four steps. Three of them are booby-trapped, and the obvious way of doing it strolls right over them like a tourist in a minefield. Here&#8217;s the damage report:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Step<\/th><th>The obvious way<\/th><th>What actually happens<\/th><th>Sanity level<\/th><\/tr><\/thead><tbody><tr><td>Install Ollama<\/td><td>Runs the official install script<\/td><td>Fails: Kaggle has no <code>zstd<\/code><\/td><td>Mildly shaken<\/td><\/tr><tr><td>Start Ollama<\/td><td>Trusts the defaults<\/td><td>4,096-token goldfish memory<\/td><td>Suspicious<\/td><\/tr><tr><td>Get the model<\/td><td><code>ollama run<\/code><\/td><td>Cell waits forever for someone to type<\/td><td>Philosophically broken<\/td><\/tr><tr><td>Open the tunnel<\/td><td>Reads 30 lines of log, then stops<\/td><td>Tunnel can freeze hours later<\/td><td>Who am I anymore?<\/td><\/tr><tr><td>Secure it<\/td><td>API key: <code>ollama<\/code><\/td><td>Protects absolutely nothing<\/td><td>Numb but enlightened<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Trap one: the installer needs a tool Kaggle doesn&#8217;t have.<\/strong> Ollama ships as a <code>.tar.zst<\/code> archive, and Kaggle&#8217;s image has no <code>zstd<\/code> to unpack it. That&#8217;s IKEA delivering your wardrobe without the Allen key. Install it first:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>!sudo apt-get install -y zstd\n!curl -fsSL https:\/\/ollama.com\/install.sh | sh<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The installer will then complain that systemd isn&#8217;t running. Ignore it. Notebooks have no service manager, so you start the server yourself, like an animal.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Trap two: the default memory is goldfish-sized.<\/strong> Ollama gives every model a 4,096-token context window unless told otherwise. Sounds like plenty, until you learn that Claude Code&#8217;s opening instructions alone are longer. And here&#8217;s the kicker: Ollama doesn&#8217;t throw an error when a prompt overflows. It quietly chops off the beginning and carries on, smiling. Your agent forgets what it was doing and you never find out why. It&#8217;s walking into a room and forgetting why you came in, except it happens on every single message. So start the server with a bigger window:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>import os, subprocess, time\n\nenv = os.environ.copy()\nenv[\"OLLAMA_CONTEXT_LENGTH\"] = \"32768\"\nenv[\"OLLAMA_KEEP_ALIVE\"] = \"-1\"\nsubprocess.Popen([\"ollama\", \"serve\"], env=env,\n                 stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)\ntime.sleep(5)<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">And yes, <code>Popen<\/code> matters. Run <code>!ollama serve<\/code> directly and the cell blocks forever, just sitting there. Staring. Judging.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Trap three: <code>run<\/code> waits for a person who isn&#8217;t there.<\/strong> The obvious way to download the model is <code>ollama run<\/code>. That command downloads, then opens an interactive chat and waits for you to type. A notebook cell has no keyboard. So it waits. And waits. Like a Tamagotchi nobody&#8217;s feeding. Use <code>pull<\/code>, which downloads and actually leaves:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>!ollama pull hf.co\/JonathanColetti\/Qwen3.8-27B-Uncensored-GGUF:Q4_K_M<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">That odd-looking name is Ollama pulling a GGUF file straight from Hugging Face. <code>Q4_K_M<\/code> is the 4-bit version, the usual trade-off between size and quality. Think of it as the model&#8217;s travel-size shampoo.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">\ud83d\ude08 What &#8220;Uncensored&#8221; Actually Means (Calm Down)<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Before anyone gets excited: &#8220;uncensored&#8221; does not mean smarter, edgier, or secretly aware of where the bodies are buried. The model was made with a technique called abliteration, using a tool named Heretic (yes, really). It finds the direction inside the model that produces refusals and removes it, with no retraining. On the author&#8217;s 100-prompt test of harmful requests, refusals fell from 98 to 12.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That&#8217;s it. That&#8217;s the whole upgrade. It is not smarter, and removing pieces of a model usually costs it a little accuracy. It&#8217;s like taking the brakes off a car and calling it a sports car. If you want a coding assistant and don&#8217;t care about refusals, the stock Qwen model is the better pick.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83d\ude87 A Tunnel Out of a Building With No Doors<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Congratulations, you now have a 27B model running on <code>127.0.0.1:11434<\/code>, inside a container, inside a Google data centre. Nothing outside can reach it. Kaggle notebooks accept no incoming connections. Your shiny new AI server is basically a genius locked in a basement.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Enter the Cloudflare quick tunnel, which solves this by going the other way. The <code>cloudflared<\/code> program inside the notebook dials <strong>out<\/strong> to Cloudflare and holds the line open. Cloudflare hands you a public address like <code>https:\/\/some-random-words.trycloudflare.com<\/code>, and requests to that address travel down the line to Ollama. No account, no domain, no firewall rules, no change request, no CAB meeting. Your GRC team would faint.<\/p>\n\n\n\n<figure class=\"szt-diagram\" style=\"margin:2em 0\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" viewBox=\"0 0 760 262\" role=\"img\" aria-label=\"Your requests reach the model on Kaggle through a Cloudflare tunnel\" style=\"width:100%;height:auto;font-family:system-ui,-apple-system,'Segoe UI',Roboto,sans-serif\" font-size=\"13\">\n<defs><marker id=\"sz-arrow\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"6\" markerHeight=\"6\" orient=\"auto-start-reverse\"><path d=\"M0 0L10 5L0 10z\" fill=\"#8a8f98\"\/><\/marker><\/defs>\n<rect x=\"0\" y=\"0\" width=\"760\" height=\"262\" rx=\"10\" fill=\"#ffffff\"\/>\n<text x=\"24\" y=\"34\" font-size=\"15\" font-weight=\"600\" fill=\"#1f2328\">Your requests reach the model on Kaggle through a Cloudflare tunnel<\/text>\n<rect x=\"342\" y=\"66\" width=\"394\" height=\"140\" rx=\"8\" fill=\"#f3f4f6\" stroke=\"#8a8f98\" stroke-width=\"1.25\"\/>\n<text x=\"358\" y=\"88\" font-size=\"11.5\" font-weight=\"600\" fill=\"#57606a\">Kaggle notebook, GPU T4 x2<\/text>\n<g fill=\"none\" stroke=\"#8a8f98\" stroke-width=\"1.25\"><path d=\"M164 146H186\" marker-end=\"url(#sz-arrow)\"\/><path d=\"M318 146H356\" marker-end=\"url(#sz-arrow)\"\/><path d=\"M468 146H482\" marker-end=\"url(#sz-arrow)\"\/><path d=\"M594 146H608\" marker-end=\"url(#sz-arrow)\"\/><\/g>\n<rect x=\"24\" y=\"102\" width=\"140\" height=\"88\" rx=\"8\" fill=\"#ffffff\" stroke=\"#8a8f98\" stroke-width=\"1.25\"\/>\n<text x=\"94\" y=\"126\" text-anchor=\"middle\" font-weight=\"600\" fill=\"#1f2328\">Your computer<\/text>\n<text x=\"94\" y=\"146\" text-anchor=\"middle\" font-size=\"11.5\" fill=\"#57606a\">Claude Code<\/text>\n<text x=\"94\" y=\"162\" text-anchor=\"middle\" font-size=\"11.5\" fill=\"#57606a\">OpenAI-style apps<\/text>\n<rect x=\"188\" y=\"102\" width=\"130\" height=\"88\" rx=\"8\" fill=\"#ffffff\" stroke=\"#8a8f98\" stroke-width=\"1.25\"\/>\n<text x=\"253\" y=\"126\" text-anchor=\"middle\" font-weight=\"600\" fill=\"#1f2328\">Cloudflare<\/text>\n<text x=\"253\" y=\"146\" text-anchor=\"middle\" font-size=\"11.5\" fill=\"#57606a\">public HTTPS URL<\/text>\n<text x=\"253\" y=\"162\" text-anchor=\"middle\" font-size=\"11.5\" fill=\"#57606a\">trycloudflare.com<\/text>\n<rect x=\"358\" y=\"102\" width=\"110\" height=\"88\" rx=\"8\" fill=\"#ffffff\" stroke=\"#8a8f98\" stroke-width=\"1.25\"\/>\n<text x=\"413\" y=\"126\" text-anchor=\"middle\" font-weight=\"600\" fill=\"#1f2328\">cloudflared<\/text>\n<text x=\"413\" y=\"146\" text-anchor=\"middle\" font-size=\"11.5\" fill=\"#57606a\">outbound only<\/text>\n<text x=\"413\" y=\"162\" text-anchor=\"middle\" font-size=\"11.5\" fill=\"#57606a\">rewrites Host<\/text>\n<rect x=\"484\" y=\"102\" width=\"110\" height=\"88\" rx=\"8\" fill=\"#ffffff\" stroke=\"#8a8f98\" stroke-width=\"1.25\" stroke-dasharray=\"5 4\"\/>\n<text x=\"539\" y=\"126\" text-anchor=\"middle\" font-weight=\"600\" fill=\"#1f2328\">Token proxy<\/text>\n<text x=\"539\" y=\"146\" text-anchor=\"middle\" font-size=\"11.5\" fill=\"#57606a\">port 8080<\/text>\n<text x=\"539\" y=\"162\" text-anchor=\"middle\" font-size=\"11.5\" fill=\"#57606a\">optional<\/text>\n<rect x=\"610\" y=\"102\" width=\"110\" height=\"88\" rx=\"8\" fill=\"#dbe7fb\" stroke=\"#2f6fde\" stroke-width=\"2\"\/>\n<text x=\"665\" y=\"126\" text-anchor=\"middle\" font-weight=\"600\" fill=\"#1f2328\">Ollama<\/text>\n<text x=\"665\" y=\"146\" text-anchor=\"middle\" font-size=\"11.5\" fill=\"#1f2328\">port 11434<\/text>\n<text x=\"665\" y=\"162\" text-anchor=\"middle\" font-size=\"11.5\" fill=\"#1f2328\">Qwen3.8 27B<\/text>\n<text x=\"665\" y=\"178\" text-anchor=\"middle\" font-size=\"11.5\" fill=\"#1f2328\">Q4_K_M<\/text>\n<text x=\"24\" y=\"234\" font-size=\"11.5\" fill=\"#57606a\">Requests run left to right and replies stream back. The tunnel dials out from Kaggle, so no inbound port is opened.<\/text>\n<\/svg><\/figure>\n\n\n\n<pre class=\"wp-block-code\"><code>import re, subprocess, threading\n\ncloudflared = subprocess.Popen(\n    [\"cloudflared\", \"tunnel\", \"--url\", \"http:\/\/127.0.0.1:11434\",\n     \"--http-host-header\", \"localhost:11434\"],\n    stdout=subprocess.PIPE, stderr=subprocess.STDOUT, text=True)\n\nfor line in cloudflared.stdout:\n    m = re.search(r\"https:\/\/[a-z0-9-]+\\.trycloudflare\\.com\", line)\n    if m:\n        print(\"Your URL:\", m.group(0))\n        break\nthreading.Thread(target=lambda: [None for _ in cloudflared.stdout], daemon=True).start()<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Two details here will save you an afternoon of swearing at your screen:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><code>--http-host-header localhost:11434<\/code> disguises outside requests as local ones. Ollama refuses anything that doesn&#8217;t claim to come from localhost, so without this flag every call comes back <code>403 Forbidden<\/code>. Ollama is the bouncer; this flag is the fake ID.<\/li>\n\n\n<li>The last line keeps reading the tunnel&#8217;s log in the background. The obvious version reads 30 lines and then stops listening, like a manager in a risk review. <code>cloudflared<\/code> keeps talking anyway, the buffer fills up, and the tunnel can freeze hours later with no error. Silent failures: the best kind.<\/li>\n<\/ul>\n\n\n\n<h3 class=\"wp-block-heading\">\ud83e\udd16 Plugging In Claude Code<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Since version 0.14, Ollama speaks Anthropic&#8217;s API as well as OpenAI&#8217;s. Bilingual. Overachiever. That means Claude Code can talk to it directly:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>export ANTHROPIC_BASE_URL=\"https:\/\/some-random-words.trycloudflare.com\"\nexport ANTHROPIC_AUTH_TOKEN=\"ollama\"\nexport ANTHROPIC_API_KEY=\"\"\nclaude --model hf.co\/JonathanColetti\/Qwen3.8-27B-Uncensored-GGUF:Q4_K_M<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Now notice what&#8217;s missing: <code>\/v1<\/code>. OpenAI-style apps such as Open WebUI or Cursor want <code>https:\/\/\u2026trycloudflare.com\/v1<\/code>. Claude Code adds its own <code>\/v1\/messages<\/code>, so if you give it <code>\/v1<\/code> too, you get <code>\/v1\/v1\/messages<\/code> and a 404. One slash. ONE. It&#8217;s the easiest mistake in the whole setup to make and the hardest to spot, which, coincidentally, describes most of IT.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You&#8217;ll also see people rename the model to <code>claude-sonnet-4-5<\/code> with <code>ollama cp<\/code>, so Claude Code finds a familiar name. It works, but passing the real name with <code>--model<\/code> does the same job without the cosplay.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">And let&#8217;s manage expectations, shall we? A 4-bit 27B model on two T4s is not a hosted Claude model in a fake moustache. It types slowly and gets lost in long, multi-step coding tasks. Great for experiments. Terrible as your daily driver. Like a free gym membership: technically it works, but we both know how this ends.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83d\udeaa The Front Door Is Wide Open (The Part Everyone Skips)<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">And now, the moment you&#8217;ve been scrolling for.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The API key in this setup is the word <code>ollama<\/code>. It looks like a password. It is not a password. It is a word. Clients refuse to start without <em>some<\/em> key, so you type a placeholder, and Ollama never checks it. Ollama has no authentication. None. Zero. Zilch. It&#8217;s a nightclub with no bouncer, no door, and a neon sign that says FREE GPUs.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So the moment that tunnel comes up, your model sits on the public internet behind a URL and absolutely nothing else. Anyone who finds that URL can:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>burn your weekly GPU hours on their own requests (thanks for the free compute, buddy);<\/li>\n\n\n<li>call <code>\/api\/delete<\/code> and wipe your model, or <code>\/api\/pull<\/code> and fill your disk;<\/li>\n\n\n<li>send whatever they like through an <em>uncensored<\/em> model, with <em>your<\/em> account attached. Enjoy explaining that one.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">&#8220;But the URL is random!&#8221; I hear you cry. Sure. Until you paste it into a screenshot, a Discord message, a shared notebook, or a log file. A random address hides the door. It doesn&#8217;t lock it. Security through obscurity: the corporate classic, now available for hobbyists.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">\ud83d\udd10 The Fix: A Tiny Proxy With a Password<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Put a small gatekeeper between the tunnel and Ollama. It checks every request for a secret token and slams the door on anything without one. Claude Code already sends <code>ANTHROPIC_AUTH_TOKEN<\/code> as a bearer token, and OpenAI-style apps send their API key the same way, so your clients need zero changes beyond swapping <code>ollama<\/code> for your secret. Look at that: actual security, no 200-page policy document required.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>%%writefile proxy.py\nimport os, aiohttp\nfrom aiohttp import web\n\nSECRET = os.environ[\"PROXY_TOKEN\"]\nUPSTREAM = \"http:\/\/127.0.0.1:11434\"\n\nasync def handle(request):\n    if (request.headers.get(\"Authorization\") != f\"Bearer {SECRET}\"\n            and request.headers.get(\"x-api-key\") != SECRET):\n        return web.Response(status=401, text=\"unauthorized\")\n    headers = {k: v for k, v in request.headers.items()\n               if k.lower() not in (\"host\", \"authorization\", \"x-api-key\", \"content-length\")}\n    async with aiohttp.ClientSession(timeout=aiohttp.ClientTimeout(total=None)) as s:\n        async with s.request(request.method, UPSTREAM + request.rel_url.path_qs,\n                             headers=headers, data=await request.read()) as up:\n            resp = web.StreamResponse(status=up.status, headers={\n                \"Content-Type\": up.headers.get(\"Content-Type\", \"application\/json\")})\n            await resp.prepare(request)\n            async for chunk in up.content.iter_any():\n                await resp.write(chunk)\n            await resp.write_eof()\n            return resp\n\napp = web.Application(client_max_size=50 * 1024**2)\napp.router.add_route(\"*\", \"\/{tail:.*}\", handle)\nweb.run_app(app, host=\"127.0.0.1\", port=8080, print=None)<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Start it with a random secret, then point the tunnel at port <code>8080<\/code> instead of <code>11434<\/code>:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>import secrets\nTOKEN = secrets.token_urlsafe(32)\nsubprocess.Popen([\"python\", \"proxy.py\"], env={**os.environ, \"PROXY_TOKEN\": TOKEN})\nprint(\"Your API key:\", TOKEN)<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Replies still stream token by token, and the secret is stripped out before anything reaches Ollama. If you own a domain, a named Cloudflare tunnel with Cloudflare Access does the same check at Cloudflare&#8217;s edge. Either way, when you&#8217;re done, close the door behind you: <code>cloudflared.terminate()<\/code>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">One more thing, before you build your entire personality around this. Kaggle gives out free GPUs for data science and ML work. I couldn&#8217;t find a rule that bans tunnels outright, but running a public model server is clearly not what the free tier is for. Read the terms before you make it part of your routine. &#8220;I didn&#8217;t read the terms&#8221; has never once worked as a defence.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83c\udfc1 Final Rant: Free GPUs, Borrowed Time<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">So, is it worth it? For the right job, absolutely.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If you want to try a 27B model before buying hardware, compare an abliterated model with its stock version, or see how far an open model gets inside Claude Code, this is the cheapest lab on the planet. Fifteen minutes and zero dollars gets you a model most laptops can&#8217;t even load.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But know what you&#8217;re holding. It&#8217;s a server that vanishes every twelve hours and changes its address every time it comes back, like a witness in protection. Its model is a step down from the hosted ones. And its front door is open until you lock it. If you need something always on, rent a GPU by the hour from RunPod or Vast.ai and get a fixed address. If you need privacy, run Ollama on your own machine.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But all sarcasm aside, the real lesson here isn&#8217;t the tunnel or the model. It&#8217;s that putting something on the internet now takes one command, and securing it still takes thought. The command is the easy part. Bring the thought yourself.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Because at the end of the day, free GPUs don&#8217;t fix stupid either.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Happy (authenticated) tunneling!<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Kaggle will hand you two free GPUs. Add Ollama and a Cloudflare tunnel and you&#8217;ve got a private AI server for Claude Code in fifteen minutes. Also a front door the size of a barn, wide open. Let&#8217;s fix that.<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_memberships_contains_paid_content":false,"footnotes":""},"categories":[7,50],"tags":[63,62,64,61,39,58],"class_list":["post-1000024","post","type-post","status-publish","format-standard","hentry","category-tech","category-tutorials","tag-cloudflare","tag-kaggle","tag-llm","tag-ollama","tag-python","tag-security"],"jetpack_sharing_enabled":true,"jetpack_featured_media_url":"","_links":{"self":[{"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/posts\/1000024","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/comments?post=1000024"}],"version-history":[{"count":12,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/posts\/1000024\/revisions"}],"predecessor-version":[{"id":1000029,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/posts\/1000024\/revisions\/1000029"}],"wp:attachment":[{"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/media?parent=1000024"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/categories?post=1000024"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/tags?post=1000024"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}