# Run experiments on your own server

> Connect an HPC cluster or remote server to Althea, and run commands and training jobs from the conversation.

The HPC daemon is a small program you run on your own machine. HPC means high-performance computing; a daemon is a program that stays running in the background. It connects Althea's [code agent](/code) to the hardware, files, and software you already use for research.

Suppose you have a training script and dataset on your lab's cluster. Ask Althea to check the environment, submit a run, and read the results when it finishes. The computation happens on the cluster, and Althea brings the job's status and output back to the conversation.

_[diagram: You ask Althea, and its code agent queues independent training runs on HPC A, the only cluster with the dataset. HPC B is idle, so the code agent copies the dataset to it through Tiptree’s cloud file store (dashed) and moves the runs that have not started (dotted).]_

## What runs where

The code agent prepares the work in its cloud sandbox. The daemon receives commands on your server and runs them under the operating-system account that started it. It opens the connection outward to Tiptree, so you do not need to open an inbound firewall port.

| Work | Where it runs |
| --- | --- |
| Quick checks, such as inspecting files or checking installed packages | A shell on your connected server |
| Training and other long jobs on a SLURM cluster | Compute nodes allocated by SLURM, the cluster's job scheduler |
| Long jobs on a server without SLURM | Background processes on that server |

Althea can continue the conversation while a job runs and receives an update when it finishes. [Weights & Biases](/stack#weights--biases) is optional experiment tracking; the daemon provides the connection to your compute.

## Connect a server

You need a Tiptree account and Python 3.10 or later on Linux or macOS. Run these commands on the cluster login node or remote server you want Althea to use:

```bash
pip install tasc-hpc-daemon
hpc-daemon setup --email you@example.com
hpc-daemon start --detach
hpc-daemon status
```

Use the email address of your Althea account. The setup interview walks you through:

1. Verify your email with the code sent to you, then read and accept the remote-execution agreement.
2. Choose built-in cluster guidance, or describe your server and review the instructions prepared from your notes. Optionally add a directory of your own Markdown skills: reference documents Althea can consult.
3. Choose the directories Althea may write to. On Linux, choose whether to require operating-system enforcement before the daemon can run.
4. Choose where jobs should run and save their files. The default is `hpc_jobs` inside the first approved directory; choosing another location outside the approved directories asks you to approve write access there too.
5. Add any server instructions, such as which partition to use or how many jobs to run at once.

After setup, `start --detach` keeps the daemon running after you disconnect; `status` checks the local process.

Directory restrictions need Bubblewrap on Linux or Apple's sandbox-exec on macOS. On a SLURM cluster, Bubblewrap must also work on the compute nodes. The [installation and safety reference](https://pypi.org/project/tasc-hpc-daemon/#user-content-requirements) covers those requirements and the command options.

## Check the connection

In Althea, open Settings from your profile menu, then choose HPC daemons under Advanced. Find your daemon ID and check its Online or Offline indicator. Click Refresh to request the latest status; the card also shows when the server was last seen and its reported capabilities.

![HPC daemons settings panel with two online example clusters and one offline evaluation server, their IDs and reported capabilities](/assets/img/case-studies/hpc-daemons-settings.png)

_The HPC daemons panel with example connections. Refresh checks their status; each card shows the daemon ID and reported capabilities._

Online confirms the platform can see the daemon. The request below checks whether Althea can run a command on it.

## Try it in the conversation

Use the daemon ID shown in settings, or run `hpc-daemon list`, then try this request in Althea, replacing `<daemon-id>` with that ID:

> On my connected server <daemon-id>, show the current directory and whether SLURM is available. Don't submit a job yet.

This asks the code agent to check the connection before doing any training. Once it works, give Althea the script and dataset paths, the resources the job needs, and where to save the results. SLURM jobs still wait for the cluster to allocate resources.

## Work across servers

For a training sweep across two connected clusters, try:

> Run these independent experiments across my two clusters, following each server's resource limits. Only the first cluster has the dataset, so copy it to the second before starting runs there.

Althea's code agent can distribute independent runs across suitable servers according to their available resources and your instructions. As jobs finish, it can reassign work that has not started when that would shorten the overall run. Experiments already running stay where they are; each cluster still allocates resources through its own scheduler.

It can also copy files or directories between your connected servers through Tiptree's cloud file store, without requiring a direct SSH connection between them. Tiptree starts automatic cleanup of the stored transfer data when the transfer ends; deletion may be delayed. Both daemons must be online, and the destination must be within its approved writable directories when restrictions are configured. Ask Althea to check transfer progress or the result.

## Your access and controls

Choose writable directories for the work you want Althea to do. These restrict writes when operating-system enforcement is active; they do not hide other files your account can read. On SLURM, a job step started with `srun` runs outside this write restriction. Running without that enforcement leaves only advisory checks. Use your own operating-system account, not one shared with other people.

Run `hpc-daemon stop` to stop the connection. This does not cancel jobs already submitted. For startup options, logs, and managing multiple daemon profiles, see the [package reference on PyPI](https://pypi.org/project/tasc-hpc-daemon/#user-content-running).
