Blog
Longer arguments about why any of this is worth doing on your own hardware. Written by the people who keep the fleet running, and opinionated where we have a view.
Private AI
What leaves the building when a team uses someone else's model, and what does not have to.
Your employees are already using AI. The only question is where the data goes.
Blocking consumer chatbots pushes them onto personal accounts, so give staff a sanctioned endpoint instead.
The best models you're not allowed to send your data to
Which open-weight models fit in 128 GB or 256 GB, and what you actually give up against a frontier API.
Agents and automation
The work that pays for the machine is usually unattended, not chat.
An AI employee needs an office. Give it one you own.
An agent with your mail, files and credentials is a staff member with root. House it accordingly.
Automate the boring 80%, with an LLM that never bills per token
On a flat-rate machine the marginal token costs nothing, so it pays to run the model over everything.
Want the hands-on version?
The docs cover the commands: SSH keys, vLLM flags, model sizing and benchmarks we ran ourselves.
Read the docs Deploy a Spark