New

Talos Linux for AI and Talos Cupar: sovereign AI on infrastructure you own. General availability March 2027.

Talos Linux for AI · Talos Cupar

Your models. Your data.
Your infrastructure.

Serve, fine-tune and ground open models on GPUs you own, without your data leaving the building. Talos Linux for AI is the open-source, immutable operating system underneath. Talos Cupar is the commercial platform that runs it.

Talos Cupar and Talos Linux for AI · General availability March 2027

Talk to us →

What you can build

Select. Train. Customize. From an open model to an assistant that knows your business.

Select

Open-weights models

Pick open-weight models from a managed library, pinned to an exact revision and verified by SHA-256. Serve them through an OpenAI-compatible API on NVIDIA or AMD GPUs you already own.

Train

SFT LoRA fine-tunes

Teach the model your language with per-tenant LoRA adapters trained on fenced GPUs. Automatic evaluation splits show before and after quality on your own held-out data.

Customize

RAG over your documents

Ground answers in your own document library with numbered citations, and connect live systems of record through MCP so agents can read and act.

How it works

The scaffolding is generic. What you bring makes it yours.

Talos provides the full path from a model to a working assistant: serving, training, retrieval and connectors. What makes the result yours are the three things only you have, your instructions, your documents and your databases. Read it from the bottom up.

Delivered through Talos Cupar

Chat

Conversational chat with memory, grounded and cited

Agentic AI

Retrieve, reason and act across your systems

Scaffolding provided by Talos

Inference harness

vLLM, OpenAI-compatible API

SFT LoRA

Training pipeline

RAG

Retrieval and vector database

MCP

Connectors and tools

You bring your customization

AI models

Open weights or frontier

Training instructions

Prompts and responses

Enterprise documents

PDF, Word, wikis

Enterprise databases

Via connectors

Declarative

Apply a config, watch it converge

GPUs, storage, model libraries and the revisions to pull are declared as documents. Apply them and the node converges, with no golden images and no shell runbooks.

Reboot-proof

Upgrades stay calm

Models re-verify in place rather than downloading again, and paired nodes roll operating system and driver upgrades one at a time with no platform downtime.

Observable

Status from one place

GPU allocation, model pulls with per-file progress, training phases and retrieval health all surface as first-class status.

Talos model training

Teach the model your voice via SFT LoRA.

01

Your data

Prompt and response pairs, in simple JSONL.

02

Validate

Format checked and an evaluation split made.

03

Train

Adapters are created for the models you selected.

04

See proof

A before and after report with real samples.

05

Go live

Served under its own name.

Talos RAG

Ground it in your documents using retrieval augmented generation.

01

Load

PDF, Word and HTML documents go into your library.

02

Ingest

Documents are parsed, chunked and embedded on GPU.

03

Ask

Hybrid vector search and reranking find the best passages.

04

Answer

The model answers from those passages, with numbered citations.

05

Verify

Every response links back to its source.

Talos MCP

Connect it to your systems using Model Context Protocol.

Agents ground answers in live enterprise data and take approved actions, such as creating a ticket or updating a record. Every call is scoped to a role and logged.

01

Connect

MCP servers for your own systems.

02

Discover tools

Agents see only approved actions.

03

Retrieve and act

Read and write systems of record.

04

Guard

Role-scoped, audited calls.

05

Perform

Workflows across systems.

Talos Cupar

One console for the AI lifecycle.

Talos Cupar is the commercial platform for running AI on private infrastructure. It turns the building blocks above into an operated service, with one place to decide what runs, on which silicon, trained on what, and who can reach it.

Select

Models and inference engines

Choose the models and serving engines that run, each pinned to an exact, verified revision.

Allocate

GPU capacity

Allocate GPUs through sharding and slicing, and repurpose capacity as demand changes.

Configure

Fine-tuning, RAG and MCP

Provision fine-tuning including LoRA, retrieval pipelines and MCP connectors from one interface.

Deliver

Chat and agentic AI

Ship conversational chat with memory and tool-calling agents, with grounded answers and numbered citations. Clients pick any base model or tenant adapter by name.

Talk to us →

Control, privacy and compliance

Sovereignty is about control, not location.

Keeping data in-country is not the same as controlling it. What settles sovereignty is who holds the keys, who can reach the machines, where models are trained and which law applies. Talos Linux for AI has a concrete answer for each.

Who holds the keys

Weights stay on your storage

Models, adapters, vectors, documents and prompts live on your hardware, in your jurisdiction. Nothing is escrowed with a provider and nothing phones home.

Who has access

Nothing to log into

No shell, no SSH, no package manager. Talos Linux exposes one mutual-TLS API, with role-based separation between operating nodes and consuming AI.

Where models are trained

On your GPUs, on your data

Fine-tuning runs on your own fenced GPUs. Every adapter carries a full lineage, from base revision and dataset hash to recipe and evaluation scores.

Tenant isolation

Clean walls between customers

Tenants are isolated at the retrieval layer and the adapter layer. One platform serves many customers without mixing their data.

Supply chain

Evidence for the auditor

Models arrive revision-pinned and hash-verified. Talos Enterprise Linux adds SBOMs, curated VEX, signed build attestations, FIPS 140-3 compliant builds, CVE SLAs and IP indemnity.

Which law applies

Runs disconnected

The whole stack runs in your jurisdiction on hardware you own, and runs air-gapped when the rules require it. Nothing meters.

The difference, side by side

A platform, from silicon to chat.

Typical GPU stack

Hand-built and hard to reproduce

  • General-purpose OS, hand-configured, drifts over time
  • Scripts download models to ad-hoc paths
  • One model per GPU budget line
  • Training and customization bolted on with a second vendor
  • Cloud API where data leaves and costs meter

Talos Linux for AI

One declarative, verifiable platform

  • Immutable OS, the whole platform in one declarative file
  • Pinned, verified, versioned model library with an API
  • Dozens of tenant fine-tunes on one shared base model
  • SFT LoRA fine-tuning with before and after evaluation reports
  • RAG retrieval, tuning and serving on the same substrate
  • MCP role-scoped, audited access to your systems of record
  • Your racks, your data, your models, with nothing metered

Availability

Talos Cupar and Talos Linux for AI.

Talos Cupar

The commercial platform for private AI

Full production support and direct engagement with Sidero Labs engineering, from onboarding onward.

General availability · March 2027

Talos Linux for AI

Free and open source, with an enterprise tier

The open-source OS is free under MPL-2.0. Talos Enterprise Linux adds the full AI suite, compliance artifacts and support under one commercial relationship.

General availability · March 2027

See it on your data

An assistant that knows your business.

Talk to us about running Talos Linux for AI and Talos Cupar on infrastructure you own, or get notified when they reach general availability.

Talk to us →

Talos Cupar and Talos Linux for AI · General availability March 2027