@drxim

· AI / LLM ops lead UTC+3

Founder of ZATVA, with 12 years in SRE and DevOps. I lead the team, keep a hand on the on-call rotation, and own architecture for client deploys. I'm also on-call myself for GPU inference and LLM serving.

vLLM GPU ops Triton CUDA Ray Kubernetes

what I do at ZATVA

GPU inference fleets and LLM serving: on-call for vLLM and Triton, plus autoscaling and cost review on prod. I hold the UTC+3 on-call shift (00:00 to 08:00 UTC), and on discovery calls I talk directly with CTOs/founders, with no sales layer in between.

What I've worked on most over the last few years is vLLM and Triton inference on A100/H100: multi-GPU autoscaling, KV-cache and batching tuning, and GPU cost optimization.

background

I started out as a classic sysadmin back in the 90s, on FreeBSD, BSDi and RedHat at an internet provider. From there it was infra and SRE on high-traffic backend services, and later a move into Web3 and AI infrastructure. Since 2024 we've worked as the ZATVA team on contract.

The fastest way to talk tech or scope a stack for your project is GitHub or Telegram via the contacts page.