Using Ollama on Carina
Ollama lets you run open-weight large language models locally on Carina’s GPUs; your prompts, code, and data never leave Stanford’s infrastructure.
Loading Ollama
Ollama is available as a module on Carina:
ml ollama
Requesting GPU Resources
Ollama runs best on a GPU. Request an interactive session via Slurm’s srun command, with an explicit --time. The normal partition defaults to a two-hour time limit, which is easy to run into once you factor in loading a model and working with it interactively:
srun --pty -p normal --gres=gpu:1 --time=04:00:00 bash
If your work is done before your session expires, release the resources with scancel so that others can use the GPU.
dev partition additionally has a two-hour maximum, regardless of what you request with --time. Use dev only for a quick test; use normal (shown above) for anything longer.See Slurm on Carina for queue options, and GPUs on Carina for guidance on when you actually need a GPU at all.
Starting the Server
Once you’re on a compute node with the module loaded, start the Ollama server:
ollama serve
Carina’s Ollama module automatically assigns a random local port each time you run ollama serve. Watch the server’s startup log for the line containing msg=Listening on 127.0.0.1:PORT. You’ll need the value of PORT for the next step.
Running Client Commands
Leave the server running in that terminal, then open a second terminal on the same compute node for client commands like ollama run or ollama list. Load the module again, and set OLLAMA_HOST=127.0.0.1:PORT.
ml ollama
export OLLAMA_HOST=127.0.0.1:PORT
ollama list
Models: Useful Commands
| Command | Function |
|---|---|
ollama list |
Show all available models on Carina |
ollama run <model-name> |
Run a model |
ollama ps |
Show currently loaded models |
ollama stop <model-name> |
Stop a running model |
ollama show <model-name> |
Display information about a model |
ollama pull, ollama create, and ollama rm will fail with a permission error, because access to the external libraries is blocked.
ollama run gemma:2b
ollama run llama3.1:8b
Batch Jobs
For longer-running or unattended work, submit Ollama as a Slurm batch job instead of an interactive session. See the Slurm Primer for sbatch basics.
Requesting a Model
If the model you need isn’t already in the shared library, request that Team Carina add it.
Request an Ollama Model
The form offers a checklist of popular models pulled daily from the Ollama Library, so you can typically just check the box next to the model you want. If a model isn’t on that list, list it (comma-separated with any others) in the form’s “Other” field. You’ll also need to provide your PI/project name. Installations are global, so a model you request will be available to everyone on Carina, not just your project.
Team Carina reviews each request before adding a model to the shared library. Once approved, the model will show up for everyone via ollama list.