Skip to main content

Using Ollama on Carina

Ollama lets you run open-weight large language models locally on Carina’s GPUs; your prompts, code, and data never leave Stanford’s infrastructure.

Loading Ollama

Ollama is available as a module on Carina:

ml ollama

Requesting GPU Resources

Ollama runs best on a GPU. Request an interactive session via Slurm’s srun command, with an explicit --time. The normal partition defaults to a two-hour time limit, which is easy to run into once you factor in loading a model and working with it interactively:

srun --pty -p normal --gres=gpu:1 --time=04:00:00 bash

If your work is done before your session expires, release the resources with scancel so that others can use the GPU.

See Slurm on Carina for queue options, and GPUs on Carina for guidance on when you actually need a GPU at all.

Starting the Server

Once you’re on a compute node with the module loaded, start the Ollama server:

ollama serve

Carina’s Ollama module automatically assigns a random local port each time you run ollama serve. Watch the server’s startup log for the line containing msg=Listening on 127.0.0.1:PORT. You’ll need the value of PORT for the next step.

Running Client Commands

Leave the server running in that terminal, then open a second terminal on the same compute node for client commands like ollama run or ollama list. Load the module again, and set OLLAMA_HOST=127.0.0.1:PORT.

ml ollama
export OLLAMA_HOST=127.0.0.1:PORT
ollama list

Models: Useful Commands

Command Function
ollama list Show all available models on Carina
ollama run <model-name> Run a model
ollama ps Show currently loaded models
ollama stop <model-name> Stop a running model
ollama show <model-name> Display information about a model

ollama run gemma:2b
ollama run llama3.1:8b

Batch Jobs

For longer-running or unattended work, submit Ollama as a Slurm batch job instead of an interactive session. See the Slurm Primer for sbatch basics.

Requesting a Model

If the model you need isn’t already in the shared library, request that Team Carina add it.

Request an Ollama Model

The form offers a checklist of popular models pulled daily from the Ollama Library, so you can typically just check the box next to the model you want. If a model isn’t on that list, list it (comma-separated with any others) in the form’s “Other” field. You’ll also need to provide your PI/project name. Installations are global, so a model you request will be available to everyone on Carina, not just your project.

Team Carina reviews each request before adding a model to the shared library. Once approved, the model will show up for everyone via ollama list.