Taught by Humans - < tbh />

This is an older piece. The information may be out of date.

Run Unlimited AI on Your Device: How Open Source AI Actually Works

By Job Faith Tumibay | 26 August 2025 · Updated September 2026

Open Source AIHugging FaceAILLM

Everything is always better when it’s free. Maybe not in AI, but it’s still free on your device!

It’s not surprising that the latest AI models need a monthly subscription. Electricity bills, developer efforts, and the machinery required to train and maintain these systems all cost money. If we want access to the results, we have to pay for them.

But if you don’t need the latest and greatest, there’s a whole world of “open source” AI models that you can download for free. Open source basically means everything that makes the AI work is public. You can download and change them - all without paying a subscription.

And while the lower cost is appealing, the bigger question is whether the model is the right fit for your use case.

For lighter tasks such as drafting blogs, summarising notes, or brainstorming ideas, a locally run model can be more than enough. For heavier tasks such as analysing consistently changing databases, working through lots of long documents, or managing large codebases - you’ll probably still want the power of a commercial AI model.

And of course, with anything free, there will always be a catch. In this case, it’s your machine. If you don’t have a computer, or if that computer is outdated or underpowered, this option might not be for you.


What Actually are Open Source Models?

When we say something is “open source,” we mean that its source code is public. Anyone can:

  • View - inspect how it works, ensuring transparency and security.
  • Modify - customise or improve it based on specific needs.
  • Distribute - share improvements based on the original software.

With AI models, there’s an extra layer to this - model weights. These are the numbers the AI learns during training. They’re not words or pre-written answers, but a huge list of numbers, often in the billions.

During training, the AI makes guesses about the next word, checks if it’s right, and tweaks these numbers to improve. After billions of tries, those numbers form patterns that let it predict the most likely next word in any situation. Without them, an AI is basically an empty brain.

A fully open source AI gives you both the code and the weights. That means you can run it on your own device, retrain it with your own data, and share it with anyone. An example is something like Mistral 7B. OpenAI’s new gpt‑oss models are similar, with the extra condition that you follow their usage policy. Mostly saying that you should follow the law.

Some companies do things differently, like Meta with LLaMa 4. Their models are technically available for free but not fully open source. It comes with a license agreement, which requires you to put “Built with Llama” on any materials you have made using Llama. Oddly, it stops being free if you have more than 700 million monthly users.


How Open Source AI Models Works

When you use a paid AI chatbot, your message is sent to that company’s servers. They run it through their own AI models, process the response, and send it back to you. All the thinking happens on their side.

With open source AI, you remove the companies in the process entirely. Your message goes straight to the model you have downloaded, and it replies on your own machine. Nothing is sent over the internet, which means better privacy since your data never leaves your device. As a bonus, it even runs offline!

Of course, there’s a balance. AI on the internet gives you convenience and it performs well on almost any device. But it needs constant connection, and your data is always being sent elsewhere. Running AI locally gives you privacy and the ability to work offline. The trade-off is that it depends on how powerful your own machine is, and setting it up can take a bit more effort.


The Requirements

Running AI locally doesn’t need a supercomputer, but the requirement will change depending on the size of the model you choose.

You’ll often see a “b” in model names, like Mistral 7b or gpt-oss:20b. This stands for billions of parameters, or the model weights. The higher the number, the smarter the model is, but this means they also require more hardware requirements.

Most of these open-source models are shared through Hugging Face, the most popular and reputable source for downloading and testing them. Think of it as the app store for AI models - except most of what’s there is free.

There are two main hardware specifications when it comes to running AI on your device. RAM and VRAM.

  • RAM is your computer’s short-term memory. This is where any open applications get processed.
  • VRAM is the graphics card version. It’s normally used to handle the visuals in games and videos - and it just so happens that the same power is also great for running AI models.

If you want to have a smooth experience, the numbers are roughly:

  • Smaller models (~7B) - run on most modern laptops with at least 16GB of RAM, 8GB if you’re patient. Think of the current MacBook Air with the M4 chip. It comes standard with a 16 GB of unified memory (Apple’s version of RAM), which is perfect for these models.
  • Medium models (13B-20B) - this is where the VRAM starts to matter, so you’ll need a graphics card with at least 8 GB. In practice, this usually means a higher-end machine.
  • Huge models (30B+) - these are best left to powerful computers, or cloud setups. They’re really for specialised use cases, like advanced research, AI development, or companies running heavy workloads. Not something you’d use for everyday tasks.

You can definitely run a model that’s a bit over your specs… but don’t be surprised if you can make a cup of tea before it replies. The upside is you’ll find out exactly what your machine can handle (and still get that tea).


Do You Need to Be a Tech Wizard?

Not really, but confidence helps. All you need is to be able to install software, follow written instructions, and move files around your computer.

Some tools, like LM Studio, make the process beginner-friendly. Others, like Ollama, use the command line which can be intimidating. There’s no need to have any programming knowledge, but if your computer skills stop at opening a browser, expect a short learning curve.


Comparison to the Top Commercial Models (as of August 2025)

OptionTypical CostsAdvantagesDisadvantages
Local (Open Source)Free on your own deviceNo subscription fees, works offline, full control over dataNeeds capable hardware, can be slower on older machines, and needs to be set up.
ChatGPT Team (GPT-5)£300/year or £30/month per memberAlways up to date, works on almost any deviceRequires subscription, needs internet
Claude Team (Sonnet)£276/year per memberStrong reasoning and performs the best on coding tasksRequires subscription, needs internet
Google Workspace Plus (Gemini Advanced)£18.40/month per memberIntegrated in Google Workspaces, more than just the AIData privacy depends on Google policies

Why Is This Important?

If all you care about is performance, then maybe paying for the fastest and shiniest AI is the way. But there are a couple of reasons why it’s worth talking about running AI locally - especially for businesses.

Personal, cost-sensitive use - if you’re experimenting or just want help with light tasks (summarising posts, drafting blogs), open-source models are a no-brainer. They’re free to run on your device, no subscription needed. You can use them as much as you want, without hitting limits.

Data privacy - if you’re reviewing contracts, analysing internal reports, or handling anything sensitive, a local model guarantees that the data never leaves your machine. Commercial models can be set up privately, but they still require trusting a third party.

Offline access - if you need AI when there’s no Wi-Fi, or in a secure facility with no internet, a local model is the only option.

Outages - AI providers go down more often than you’d think. When an online service is unavailable, you’re stuck waiting. When your local model “goes down”, it usually means your laptop is closed.

Having said that, local models come with trade-offs too.

Performance and speed - if you’re building something customer-facing where every second counts, commercial AI usually wins. Unless you have a strong engineering team and the right hardware, renting a commercial model can be simpler. Sometimes, they’re even cheaper once server costs are factored in.

It’s less about replacing paid AI with a local one - but knowing when it makes sense to keep the work in-house, and when it’s worth sending it to the cloud.