LLMOps Engineer
Take language models into production and run the serving, inference and observability behind them.
About the role
You take customer language models into production on our GPUs or theirs. You design serving stacks that are fast, stable and cost-controlled, then run them after launch.
Responsibilities
- Deploy open-weight models on vLLM, TensorRT-LLM or Triton.
- Tune batching, quantisation and caching to hit latency and throughput targets.
- Build pipelines for RAG, evaluation and model releases.
- Set up API gateways with authentication, rate limits and usage tracking.
- Build observability for latency, errors, token usage and cost per request.
- Plan capacity with the GPU infrastructure team.
Requirements
- 3+ years in DevOps, platform or MLOps engineering.
- Fluent with Kubernetes, Helm and containers.
- Strong Python; comfortable reading inference-server code.
- Has deployed LLMs or deep learning models to production.
- Experience with Prometheus, Grafana or OpenTelemetry.
Nice to have
- Understanding of KV cache, speculative decoding or tensor parallelism.
- Experience with vector databases such as Milvus, Qdrant or pgvector.
- Open source work on LLM tooling.
Tools you will use
- Kubernetes
- Helm
- vLLM
- TensorRT-LLM
- Triton
- Python
- Prometheus
- OpenTelemetry
Working at Zotect
- Work across the full AI stack, from data centres, networks and GPUs to model serving.
- Your work runs in production for customers in Thailand, Singapore and the USA.
- Work closely with the CEO in an engineering-led team.
- Our office is in the Gaysorn Building on Ploenchit Road, next to BTS Chit Lom.
How to apply
Send your CV and any portfolio links to hr@zotect.com with the role in the subject line.
Send only what we need to assess your application. Leave out ID card numbers, religion and health information. See our Privacy Policy
