← Back to Guides
AITutorialLocal LLMsHugging Face

A Complete Guide to Using Local AI Models for Everyday Tasks

Skyline Apps ·

Most AI tools you’ve used - ChatGPT, Claude, Gemini - run on someone else’s server. You send them your data, they send back an answer, and somewhere a meter is running. For a lot of everyday tasks, that’s overkill. You don’t need a frontier model to write a cover letter draft or sort your reviews into categories, and you definitely don’t need to send that data to a company’s server to do it.

This guide walks through running AI models entirely on your own computer - for free, privately, forever - and picking the right model for the task instead of reaching for the same one every time.

Why bother running models locally?

  • Privacy - your data never leaves your machine. This matters more than people think: a resume, a journal entry, a job posting you haven’t told your employer about.
  • Cost - once it’s set up, it’s free. No per-request billing, no rate limits, no surprise invoice.
  • Availability - works offline, works forever, doesn’t disappear if a company changes its pricing or shuts down a product.

The tradeoff: local models are usually smaller and less capable than the biggest hosted models. For a lot of everyday tasks - drafting text, classifying things, simple Q&A - that gap doesn’t matter.

Two ecosystems, and when to use each

This trips people up, so it’s worth being explicit: there are two different things people mean by “run a model locally,” and they solve different problems.

Ollama - for chat-style, text-generation tasks

Ollama is a tool that runs conversational/text-generation models locally with a simple interface. If the task is “write something,” “answer a question,” “have a back-and-forth conversation” - Ollama is the right tool.

Setup:

1. Download Ollama from ollama.com/download
2. Open a terminal and run: ollama pull gemma3:12b
3. Chat with it directly: ollama run gemma3:12b

That’s it. gemma3:12b is a good general-purpose starting model - capable enough for most writing/summarizing tasks, small enough to run comfortably on a modern laptop.

You can also call it from a script instead of chatting interactively:

import json, urllib.request

def ask_ollama(prompt, model="gemma3:12b"):
    data = json.dumps({"model": model, "prompt": prompt, "stream": False}).encode()
    req = urllib.request.Request(
        "http://localhost:11434/api/generate",
        data=data,
        headers={"Content-Type": "application/json"}
    )
    with urllib.request.urlopen(req, timeout=120) as resp:
        return json.loads(resp.read())["response"]

print(ask_ollama("Summarize this in one sentence: ..."))

No API key. No account. This is a plain HTTP call to a server running on your own machine.

Hugging Face - for task-specific models

Hugging Face hosts hundreds of thousands of models, but most of them aren’t chat models - they’re built for one specific job: classifying text, detecting sentiment, captioning images, transcribing audio. If your task isn’t “have a conversation” but “label this,” “score this,” “extract this” - you want a Hugging Face model via the transformers Python library, not Ollama.

Finding the right model is the part people get wrong. The homepage sorts by “Trending,” which surfaces huge, bleeding-edge research models that are irrelevant for a simple task. Instead:

  1. Click the Task you actually need in the sidebar (Text Classification, Image-to-Text, etc.) - not a vague keyword search. Model names don’t always contain the obvious keyword; a well-known sentiment model is literally named after the dataset it was trained on, with no mention of “sentiment” anywhere in its name.
  2. Drag the Parameters slider down (e.g. under 1B) to filter out giant models you can’t run anyway.
  3. Sort by Most downloads, not Trending. Downloads are a better trust signal than recency.
  4. Open the model card and use the pipeline() code snippet it provides - almost every model page has one.
from transformers import pipeline

classifier = pipeline("sentiment-analysis", model="nlptown/bert-base-multilingual-uncased-sentiment")
print(classifier("This app is great but way too expensive"))

pipeline() handles batches too - just pass a list instead of one string, with a batch_size argument for efficiency. You don’t need to drop into manual PyTorch unless you need very specific control over the model’s internals or you’re operating at a scale where every millisecond matters - neither applies to a personal daily-life tool.

A full worked example

Here’s a real one: a tool that generates a tailored cover letter from your resume and a job posting - entirely offline via Ollama.

Step 1 - write the prompt. The first version looked like this:

prompt = f"""
Their background: {resume_text}
The job posting: {job_text}

Write a cover letter that opens with genuine interest in this specific
role/company (infer company/role name from the posting if present)...
"""

Step 2 - test it, and catch what actually happens. In testing, when no company name existed in the posting, the model didn’t skip mentioning it - it left a literal [Company Name/Role Name - if known, otherwise omit] placeholder bracket sitting in the output. This is a common failure mode: ambiguous instructions get followed literally, not sensibly. Telling a model to “infer if present” doesn’t tell it what to do when it’s not present.

Step 3 - fix the actual ambiguity, not just add more words.

prompt = f"""
Their background: {resume_text}
The job posting: {job_text}

Write a cover letter that opens with a plain greeting ("Dear Hiring
Manager," unless a specific name/company is clearly stated) followed by
genuine interest in the role.

Never use bracketed placeholders like [Company Name] - if you don't know
a detail, just don't mention it.
"""

One sentence removed the ambiguity entirely, and the bug disappeared on re-test. This is the actual rhythm of working with these models: write something reasonable, run it on a real example, look closely at what came back (not just whether it “looks okay”), and fix the specific gap you find - not a vague “make it better” pass.

Step 4 - package it so you’ll actually use it again. A model call buried in a notebook doesn’t get used twice. Wrap it in a small script with clear inputs and outputs:

import argparse

parser = argparse.ArgumentParser()
parser.add_argument("--resume", required=True)
parser.add_argument("--job", required=True)
args = parser.parse_args()

resume_text = open(args.resume).read()
job_text = open(args.job).read()
# ... build prompt, call ask_ollama(), print + save result

Now it’s python3 my_tool.py --resume resume.txt --job job.txt instead of re-copying code into a notebook every time.

Where this actually goes

The same pattern - pick the right ecosystem, find the model with the Tasks filter (not Trending), write a prompt, test it on a real example, fix the specific gap you find, wrap it in a script - applies to almost anything: summarizing your own notes, sorting messages, drafting replies, extracting structured info from a document. The model choice changes; the process doesn’t.

If you want to see this pattern applied end-to-end, PitchCraft - the free tool this example was drawn from - is a working, downloadable version of exactly what’s described above.