10 Security Hacks Every Local AI User Should Know

You can run AI models locally without risking rogue cyberattacks.

10 Security Hacks Every Local AI User Should Know

When you think of local AI, you might assume that running models on local hardware instead of paying for cloud-based services is the safer and more private option. That’s only partially true. When running AI locally, you maintain better control over your workflows and data. You decide whether to share your data with any third-party services, and if so, under what terms. You also don’t have to worry about data breaches and cyberattacks targeting major tech corporations. But on the other hand, you refuse the multimillion-dollar security infrastructure built by established players like OpenAI or Anthropic. Your security is now fully your responsibility, for better or worse. With that in mind, here are ten clever hacks to help you achieve better security when running AI models on your PC or a private VPS (virtual private server). 

Is it safer to run AI models on local hardware?

When running AI models on your personal hardware using a platform like Jan, Ollama, or LM Studio, your messages, documents, and chat history do not leave your device and aren’t sent to someone else’s cloud servers. If you’re worried about an AI company selling your data without your consent, or if you don’t want to end up with your credentials leaked in the next data breach, going local is the smart move. 

That said, it isn’t foolproof. You’re still downloading the AI model files from a public database. You may also have to let the model access certain APIs so it can talk to your software or data over the public web. Finally, if you’re using a public wifi, anyone else connected to the network may be able to breach your operating system by targeting the AI. In other words, things are not as simple as people sometimes make them out to be. 

Just this January, SentinelOne and Censys found 175,000 publicly exposed Ollama hosts that could be used by any attacker with an internet connection to execute code and connect to third-party services from a user’s credentials and hardware. If you want to run AI locally, you have to be very careful with where you get your models from and what they have access to. Here are some tips to help you get it right.

Keep your model server on localhost

To run AI models on your local machine, you’ll need to use an inference engine (also called a runner) like Ollama or LM Studio, which let the model load and execute on your hardware. By default, AI runners are configured to run models on localhost (127.0.0.1 or 0:0:0:0:0:0:0:1), meaning that other devices on your network or the public web can’t access it. But if you run your AI model on 0.0.0.0, that opens up access to all devices in your network. Anyone on your shared wifi can boot up your local AI setup, then use it to make changes to your hardware or steal sensitive data.

Sometimes, setup guides will suggest that you do this anyway, so that you can access your local AI model from other devices on your network, like a smartphone or laptop. It also comes up when people try to run AI models on VPS servers or Network-Attached Storage (NAS) devices. But this will put your data and workflows at risk, so if you did something to change the default server configuration of your model runner, make sure to change it back now: 

On Ollama, you can do this by changing the OLLAMA_HOST variable back to 127.0.0.1. 

For LM Studio, toggle off “Serve on Local Network."

If you use Jan, click the gear icon on your Hub interface to get to the Settings page. Then select Local API Server and make up an API key using an online generator like RandomKeygen.

Use a private VPN tunnel instead of port forwarding

You shouldn’t expose your AI model to your public IP address on the internet. But what if you still need to share model access to your other devices remotely? Normally, people enable port forwarding on their routers to configure access to their resources and data from a remote location. But you should never use this approach to configure remote access to AI models or runners on your local machine. 

If you’re already running your AI model on 0.0.0.0, and you choose to enable port forwarding on your router on top of that, anyone on the internet can break into your local AI setup if they manage to guess your IP address. Cyber attackers often operate bot networks that routinely scan the internet for open ports on residential IPs, so you're running the risk of being targeted if you do this. A better way is to set up an encrypted tunnel using a VPN or Cloudflare ZTNA. Mesh VPNs like Tailscale are a popular choice for this, as is Cloudflare Zero Trust’s new Tunnel feature. 

Update your AI runner as soon as patches land

In May 2026, Cyera uncovered a new Ollama vulnerability that let attackers steal chunks of your data and credentials using unauthenticated API calls. The flaw, called “Bleeding Llama,” had a CVSS rating of 9.3 out of 10. At the time, it put around 300,000 publicly exposed Ollama servers at risk until it was addressed in patch version 0.17.1. 

AI runners like LM Studio, Ollama, Jan, and GPT4All are still experimental and often reveal new vulnerabilities that get patched in subsequent releases. If your runner is even a few versions out of date, your server could be vulnerable to a serious attack vector that hackers can exploit. Always grab the latest release as soon as you can from the AI runner’s official website or GitHub repository.

Choose safetensors or GGUF files over pickle

AI models based on older deep learning models like PyTorch are often downloadable as pickle files, with extensions like .bin, .pt, or .pkl. But due to the nature of the Python pickle file format, these model files can be altered to execute malicious code as soon as you try to load them using your runner. 

Back in 2025, ReversingLabs found two live model files on Hugging Face that had cleared the platform’s automated security checks even though they had an unauthorized remote access function hidden in plain sight. Now, Hugging Face’s own documentation notes pickle files as a major security risk.  

To avoid data breaches or unauthorized access, you should only download LLMs that come packaged in newer file formats like .safetensor or .gguf. These file formats store your data in numerical format, which makes malicious code execution impossible as a model loads. If a particular model is only available as a .pt  or .pkl file, I’d just skip it. There are plenty of newer-version LLMs that use more secure file formats. 

Download models from publishers you can verify

AI hubs like Hugging Face or ModelScope allow anyone with an internet connection to upload AI models to their website. While they have some platform-level security protocols in place, in cases like the incident discovered by ReversingLabs in 2025, newer or more sophisticated exploits can bypass these protocols and verifications very easily. 

For better safety, download model files uploaded from official accounts managed by major model developers only. For example, Google, Mistral, Meta, and Qwen (Alibaba) all have separate organizational accounts with a verified badge on Hugging Face.  Verification badges indicate that a company account is really owned and administered by that company, because the uploader would have had to use an official company email address to log in and upload the model files. You can see the Advanced Security section of Hugging Face’s documentation for more details on how verified badges work for enterprises.

Get your AI apps from official websites only

Hackers like to use popular GenAI tools as a lure to get people to install malicious software. Often, they’ll set up fake websites or upload to popular app marketplaces where they can pose as official platforms. Earlier attempts focused on ChatGPT clones on lookalike websites that seemed like the real OpenAI. They would get people to download a corrupt .exe or .dmg file, which would then deliver dangerous payloads like Redline, Lumma, or the Odyssey infostealer for Mac. Similar attempts have also been used to target Android users through malicious apps uploaded to the Play Store. 

What do you think so far?

But the attacks have grown more sophisticated since then and may even target obscure local AI platforms and Python packages. Attackers have gone as far as to breach official GitHub repositories and Python Package Index (PyPI) uploads. TrendAI reported one particularly disturbing instance where malicious code was inserted directly into the official PyPI package of LiteLLM, an open-source AI gateway that lets you call hundreds of LLMs from a single API. Positive Security also discovered malicious Python packages uploaded to PyPI as Deepseek lookalikes. 

Make sure to verify where you’re getting your AI tools from. Your best bet is to rely on direct official sources, verified GitHub repos maintained by trusted AI vendors, and Python packages referenced directly in the source company’s official documentation.

Double-check packages your model tells you to install

I already covered how Python packages are corrupted to install malware as soon as you run them on your system. But it’s not just the LLM files and AI tools that you need to watch out for. When you ask AI agents to write code or execute tasks, they also install and run any packages or dependencies needed to complete that job. And because AI models are prone to hallucination, agents will often just make up package names that don’t exist. A recent study that analyzed 16 models across 576,000 code samples found that open-weight LLMs do this 21.7% of the time, while frontier AI models have a lower hallucination rate of 5.2%. 

Hackers know this, hence "slopsquatting," a new attack in which bad actors register fake software packages under commonly hallucinated package names across different LLMs. These packages can run malicious code, prompt injection attacks, or infostealers as soon as your AI agent runs them on your local machine. 

The best way to avoid these attacks is to limit what your AI agent can install and run without your approval. You can either choose to manually approve each software package before the model installs or runs it, or you can whitelist certain trustworthy repositories that aren’t likely to contain malware. Either way, make sure to review your model’s log to see what pip install and npm install commands it runs to avoid unauthorized installations.

Limit what your AI agents can touch

Even when you run them on your local hardware, AI agents can call MCP servers, download and run files, search the web, or connect to third-party services using APIs. Moreover, they can read and write files to your local hardware and even change core operating system settings. All of these features should be enabled only with an abundance of caution based on your security profile. Carefully manage the level of access an AI agent or model runner has on your system, especially newer open-weight models that are more likely to hallucinate or have exploitable vulnerabilities. 

There are multiple ways to regulate how much access an AI agent has. The first is to run your AI workflows inside a Dockerized container that can’t make direct changes to your system files. Beyond that, you can also restrict permissions by changing the default configuration of your agentic framework, like OpenClaw or Hermes. OpenClaw lets you choose between three default permission profiles, including ask, deny, and allowlist, which can be further scoped to specific workflows and services. Hermes also lets you set up a similar allowlist (whitelist) or restrict tool usage per cron job. 

Switch on local-only mode

AI model runners like Ollama and LM Studio can support local as well as cloud-hosted models. But you can configure them to restrict network access through a single function even if you haven’t done so at the orchestration layer with Hermes or OpenClaw. You do this by binding the service to your local IP address (127.0.0.1) to prevent other devices from accessing it over your network or the public web.

Encrypt the drive that holds your chat history

When you keep your AI workflows local, your entire chat history, along with any credentials, secrets, or API tokens you may have shared with your model, exist in plain text on your local drives. If someone managed to access your device physically, they could take all of it. Apps like FileVault, BitKocker, or LUKS can encrypt your hard drive so that your chat history can’t be read in plain text without an encryption key to decode it. Use them to avoid the risk of exposure if your device is stolen or lost.

A few local AI platforms to get started with

If you’re new to local AI, here are a few platforms to play around with. They offer the best accessibility for new users who aren’t familiar with the technicalities of AI engineering. 

Ollama: An open-source model runner for macOS, Windows, and Linux. Large model library and a simple desktop app that most other local AI tools can plug into.

LM Studio: A polished desktop app that lets you download models from Hugging Face inside a graphical UI. It's been free for both work and personal use since July 2025.

Jan: An open-source, Apache 2.0-licensed ChatGPT alternative that runs fully offline on Windows, macOS, and Linux.

AnythingLLM Desktop: A free MIT-licensed app for chatting with your own documents locally. It's a solid pick if you want to feed PDFs and notes to a model without uploading them anywhere.

Open WebUI: A browser-based offline chat interface that can connect to Ollama and make the UI more accessible. Pair it with a mesh VPN, and your whole household can use one AI server safely.