Private AI Infrastructure · halo offering

Own your AI outright. No big tech company in the loop.

For the shops that legally or reputationally can't send client data to a public cloud API — legal, medical, financial. We build private AI on hardware you own, or on a GPU cluster billed to your own cloud account. Open models, running in your building or your account. Your data never routes through OpenAI, Google, Anthropic, or anyone else.

Client-owned hardware — the title's in your nameOpen models (Kimi, DeepSeek, GLM, gpt-oss) on-premData never leaves your building or your cloud accountNo per-token bill to a third party
Why own it instead of renting an API

Data sovereignty, priced honestly.

It never leaves the building

The model runs on hardware you own, on your network. Privileged files, patient records, financial data — none of it is uploaded to a public API to answer a question. That's not a setting you trust; it's a wire that isn't connected.

You own the hardware and the title

Same doctrine as everything else we build: it's registered to you. Hardware is invoiced as its own line at a 10–15% sourcing markup — not a 3× agency markup — because this crowd checks Nvidia's MSRP and we'd rather show the receipt.

No per-token meter

Public-cloud AI bills you every request forever. Owned hardware is a capital cost, then electricity. For a shop running heavy volume on private data, the math flips toward ownership fast.

Option A · client owns the hardware

Three sizes, from a single node to a four-way cluster.

Turnkey totals below are hardware pass-through (at a 10–15% sourcing markup, disclosed as its own line) plus a fixed integration fee. Hardware is invoiced and collected in full before anything is ordered; the integration fee follows the normal half-up-front, half-on-delivery.

Starter

Proven
$5,000–7,000
typical turnkey total

A single node that sits in your office and runs a capable open model with nothing leaving the box. Proven — Kris has run this exact single-node FP4 config personally (gpt-oss:120b at 30–40 tokens/sec).

Hardware
1× AMD Ryzen AI Max+ 395 mini-PC, or 1× DGX Spark
Model class
30B-class open model (Qwen3.6-27B, gpt-oss-120b)
Hardware (pass-through +10–15%)
$1,500–4,000
Integration fee
$2,000–3,000

Pro

Founding-client
$14,000–27,000
typical turnkey total

A small clustered build for frontier-class open models and long context. Founding-client territory: the first clustered builds disclose openly that cross-node throughput at scale is still being characterized.

Hardware
2–3× DGX Spark cluster
Model class
Kimi K2.7 / DeepSeek V4-Flash, 250K+ context
Hardware (pass-through +10–15%)
$8,000–19,000
Integration fee
$4,000–8,000

Enterprise

Founding-client
$25,000–35,000+
typical turnkey total

A full private stack wired into your systems, with retrieval over your own documents. Also founding-client territory — scoped and priced per milestone, with the same honest caveat on clustered throughput.

Hardware
4× DGX Spark cluster, integrated into client systems
Model class
Full agentic setup, custom RAG on your own documents
Hardware (pass-through +10–15%)
$16,000–19,000
Integration fee
$8,000+ across scoped milestones

Straight talk on maturity: Starter is proven — Kris has run that exact single-node config personally. Pro and Enterprise clustered builds are offered as founding-client work: cross-node throughput at scale is still being characterized, and the first clients hear that up front rather than reading it in a footnote.

Option B · you own no hardware

Managed private cloud / GPU cluster.

Want dedicated GPU compute — say a stacked multi-H100 cluster — without buying and housing hardware? We architect and manage it on a cloud GPU provider. Kris has built these before, including H100 clusters on Northflank. You hold the cloud-provider billing account directly, so the studio never fronts or is liable for GPU-hour spend — it's your account, your cluster.

One-time

Cloud AI Cluster Setup

$4,000–8,000fixed, per scoped milestone

Provider selection, cluster architecture, model deployment, and integration. Provider choice is part of the deliverable, not a fixed default — Lambda Labs for reliability first, RunPod for budget/spot-tolerant workloads, CoreWeave for enterprise scale needing InfiniBand, Northflank when you want GPU plus app hosting on one platform.

Recurring

Cloud AI Ops

$500–1,500/moscaled to cluster size

Ongoing cluster management, scaling, monitoring, and cost optimization. Your GPU-hour bill is separate and paid direct to the provider — we manage the cluster, you own the account.

Keep it healthy

AI Infrastructure Care.

AI Infrastructure Care

$250–750/moscaled to tier

Model updates, tuning, monitoring, and troubleshooting on your client-owned Private AI hardware — the local mechanic keeping your private stack current instead of letting it rot after handoff. Optional, cancel anytime.

Talk through a plan
No email required

If your data can't go to the cloud, let's talk.

This is a hands-on build, not an off-the-shelf product — every engagement starts with a call to scope your data, your compliance constraints, and the right size. Grab a time on my calendar, or head back and use the contact form on the main page; it goes straight to my phone.

← Back to Kriisko Studios