Distributed AI infrastructure

Distributed AI Compute,
Powered by Everyone.

Horizon turns compatible machines into a coordinated compute network for AI inference. One control plane for your hardware, models, and workloads.

Built for trusted, team-controlled workers
Deployment path HZN / 01
DeveloperAPI request
Control PlaneRoute · observe
Connected computeWorker status
macOS worker
Windows worker
Linux worker
Model runtimeOllama · qwen2.5:3b
Inference ready
The control plane

Operate compute,
not a collection of machines.

See workers, deployments, model runtimes, and health in one focused workspace. The preview below uses demonstration configuration based on Horizon's current MVP concepts.

Overview
Deployments
Hardware
Models
Developer
Workspace / OverviewMP
Workspace snapshot · Demo configuration

Compute overview

Connected resources and active workloads.

All systems operational
Workers3 connected
Deployments2 active
Available memory42.8 GB
Active deployment
qwen2.5:3bOllama · windows-01
Running
/v1/deployments/qwen-production/inference
Hardware
windows-01RTX 3050 · 4 GB
macbook-proApple GPU
render-node-03Offline
What is Horizon?

A control layer for fragmented compute.

AI developers often have hardware in several places: personal workstations, lab machines, development PCs, and unused CPU or GPU capacity. Horizon gives that hardware one operational boundary for inference.

01Compatible machines

Different operating systems, architectures, and resource profiles.

02One control plane

Register hardware, target deployments, and observe runtime state.

03Trusted by design

The current MVP is intended for machines you own or explicitly control.

The problem

AI workloads need somewhere to run.

Cloud GPU infrastructure can be expensive. Personal hardware is often underused. Heterogeneous machines are difficult to operate consistently.

01

Compute is fragmented

Resources live across laptops, workstations, and lab machines with different capabilities.

02

Operations are scattered

Model runtime state, hardware health, and deployment configuration should not live in separate scripts.

03

Cloud is not always the answer

For trusted teams and local workflows, existing hardware can be the practical starting point.

How it works

A small, explicit path from machine to inference.

01

Connect compute

Install and run the Horizon Worker Agent on a compatible machine.

02

Register hardware

The worker reports hardware and availability to the Control Plane.

03

Deploy a model

Choose a model and target compatible compute for the workload.

04

Run inference

The model runtime serves requests on the selected worker.

05

Observe state

Track worker status, deployment state, and runtime health.

Compute network

Different machines.
One operational view.

Horizon is designed around heterogeneous, trusted compute. Hardware stays where it is; the Control Plane provides the shared language for managing it.

Resource-aware worker metadata
One deployment targets one worker
Runtime health before serving
Control Plane central state
MmacOS workerArchitecture-aware
WWindows workerGPU-aware
LLinux workerCapability-reported
Worker connection Runtime state
Developer workflow

Build against infrastructure.

Horizon is built to make the control plane the interface. Configure workloads, inspect deployments, and integrate through clear HTTP APIs.

Read the developer docs
request.json
GET /api/models

POST /api/deployments
{
  "modelId": "...",
  "workerId": "..."
}
Heterogeneous by nature

Your machines do not need to look alike.

The current MVP focuses on trusted, team-controlled hardware. Support is shaped around what the worker can report and what the runtime can use.

macOSApple SiliconArchitecture-aware resources
WindowsNVIDIA GPUGPU and VRAM metadata
LinuxCPU / GPUPlatform support evolves with the agent
Built with boundaries

Infrastructure is also about what stays separate.

Workers do not touch PostgreSQL

Worker communication goes through the Control Plane. Database access remains a control-plane concern.

Structured commands

The system is designed around explicit deployment and runtime messages rather than arbitrary shell access.

Metadata stays distinct

Model metadata and configuration are separate from model weights and local runtime state.

Where it fits

A practical starting point for local infrastructure.

Personal infrastructure

Use available hardware for local model serving and development workflows.

Research

Coordinate compatible machines across a research environment.

Developer labs

Pool trusted hardware without immediately renting more cloud capacity.

Small teams

Expose internal compute through one operational control plane.

System architecture

Keep the boundaries clear.

Workers communicate with the Control Plane. PostgreSQL remains behind it. The deployment path is explicit from request to runtime.

DeveloperHTTPS request
Horizon Control PlaneRouting · health · state
HTTPS / WebSocket
Database access stays here
PostgreSQLMetadata and sessions
Connected workersOutbound connections
Model runtimeOllama · one worker per deployment
Inference response
Start with the machines you have

Your machines are compute.

Connect compatible hardware, deploy models, and build on top of a distributed AI runtime.