> embedded ops · devops · sre · dataops

Your DevOps team on contract. For Web3, AI, ZK and DePIN.

We know where to source hardware fast and how to ship nodes fast. We run your infra, so you don't hire SREs for four months only to lose them in eight.

team@zatva:~$
devops status --client=$YOU
@drxim UTC+3 vLLM · GPU ops @shell UTC+0 k8s · validators @serg UTC-5 provers · ZK next handoff in --:--  ·  147 runbooks · 19d since last Sev-1

pricing

We're a team of senior DevOps engineers, so we work off a single rate. The "starting at" prices below are guidance — the full grid comes with your 48h plan.

[ 48h plan ]

$1,500

Deployment plan in 48h. Fully credited toward your first contract.

[ Deploy ]

from $6k

Validators, GPU clusters, testnet bursts. One-time, and delivered in days.

[ Operate ]

from $4k/mo

24/7 on-call, signed SLA, post-mortem after every Sev-1. Monthly.

who it's for

Web3 / Validators · RPC · Testnets

Planning a testnet launch next quarter? Still can't find an SRE who actually knows Cosmos SDK after months of looking, and the one you found wants equity on top? We get your validators live in 3 regions in 5 days, with slashing alerts and a signed uptime SLA.

AI / LLM Inference · Fine-tuning

Closed your round, bought the GPUs, and now the inference bill is bleeding faster than the product ships? Your ML team can train but doesn't want to get up at 3 AM for vLLM, OOMs and Triton? We run the inference layer (autoscaling, tracing, cost-per-token), so your researchers stop being on-call.

ZK / Prover farms

Proof generation is GPU-bound, deadline-bound and parallel all at once, so one dropped job and you miss the block. We build the farm, the queue, the retries, and per-circuit benchmarks on SP1 / RISC Zero / Boundless / Brevis.

DePIN / Distributed networks

The network pays for uptime, not excuses. Running 500 nodes across 10 regions by hand has turned into a part-time job no one on your team signed up for? We onboard, monitor and rebalance the fleet, and reconcile rewards weekly. We work with Filecoin, Akash, io.net, Render, Gensyn.

how we engage

01
Discovery
1 call. Scope, stack, regions, deadlines, SLA targets.
02
Plan
Deployment plan in 48h. Architecture, hardware selection, milestones, budget.
03
Deploy
Delivery via Terraform. Repo + IaC + monitoring + runbooks as one package.
04
Operate
Signed SLA. 24/7 on-call. Post-mortem after every Sev-1.

what we run for ourselves

Map of the ZATVA fleet

We run our own fleet of 50+ nodes across 5+ countries, at 99.982% uptime. It's our proving ground: we break in the load, dashboards, runbooks and on-call rotations on infra we pay for ourselves, before any of it reaches you.

[ See the fleet → ]

cases

ZK rollup · 6 mo · validator ops + prover farm

slashing: 0 · downtime: 11 min/90d

LLM startup · 4 mo · vLLM cluster in 3 regions

cost/token: −60% · p95 latency: 380ms

DePIN sub-operator · ongoing · 200 nodes in 8 regions

uptime: 99.94% · reward tier: top-10%

Incentivized testnet · 8 wk · 50 nodes burst

top-5 operator · onboarding in 72h

stack we operate

Web3: Cosmos SDK Geth Reth OP Stack Arbitrum Orbit Polygon CDK EigenDA Celestia
AI / LLM: vLLM Triton TensorRT-LLM NVIDIA H100 / A100 Ray Kubeflow
ZK: SP1 RISC Zero Boundless Brevis Jolt Halo2
DePIN: Filecoin Akash Render io.net Gensyn
Platform: Kubernetes Terraform Ansible Prometheus Grafana Loki OpenTelemetry PagerDuty

FAQ

No, we're a managed DevOps team. We deploy and operate infrastructure on the cloud or bare metal you own, and if you don't have your own hardware, we'll help you pick it and point you to where to rent. So you pay for the team, not the nodes.

A single L1/L2 validator runs 5-10 business days end-to-end. A GPU inference cluster across 3 regions takes 10-14 days. And a burst of up to 100 nodes for an incentivized testnet we turn around in 72h.

A mix: tier-1 clouds (AWS / GCP / Azure / Hetzner / OVH / Latitude), bare-metal partners (Latitude.sh, OpenMetal), and regional providers in 12+ countries. For each project we pick by latency, price and supply window.

Retainer plus project work in cash, and tokens are optional on a case-by-case basis. Equity-only engagements are the one thing we don't do.

You do. We work through an HSM/KMS workflow where keys never leave your control: we sign, we don't custody.

It's tiered. By default: p95 first response 15 min for Sev-1 and 1h for Sev-2, and higher tiers come with dedicated on-call.

Kubernetes-first, Terraform for IaC, Prometheus + Grafana + Loki for observability, PagerDuty for on-call. That said, we adapt to the client's environment.

The deployment plan has a fixed turnaround (48h), and the price depends on scope: send a short brief and we respond within 24h. Beyond that, it's hourly or monthly.

Often yes — it depends on region and GPU type. Send the spec and we'll quote the supply window within 24h.

NDA by default, and in public cases we anonymize the details.