Cognitive Infrastructure Engineer
Optimize high-throughput LLM serving infrastructure, tensor parallel clusters, vLLM / TensorRT-LLM runtimes, dynamic batching, and hardware-accelerated model routing.
Role Overview & Operational Scope
AI models are only as good as the infrastructure serving them. You will ensure Cehpoint's proprietary models and agent clusters run with rock-solid reliability, sub-15ms time-to-first-token, and peak GPU utilization across cloud and on-premise hardware deployments.
Key Responsibilities & Production Deliverables
- Deploy, scale, and optimize inference engines (vLLM, TensorRT-LLM, TGI) across multi-GPU NVIDIA clusters.
- Implement PagedAttention, continuous batching, and speculative decoding to double token generation throughput.
- Configure automated model caching, weights distribution, and failover across regional edge locations.
- Design comprehensive Prometheus/Grafana monitoring dashboards tracking KV cache usage, queue depths, and token latency.
- Maintain rock-solid security hardening across cloud VPS instances, container runtimes, and VPC networks.
Mandatory Foundational Knowledge
- GPU microarchitecture, CUDA execution hierarchy, high-bandwidth memory (HBM), and NVLink interconnects.
- KV-cache mechanics, tensor parallelism, pipeline parallelism, and weight quantization techniques.
- Linux kernel networking, container isolation, and enterprise security compliance.
Mandatory Practical Skills & Architecture
- Expertise in Linux system administration, bash automation, and infrastructure-as-code.
- Deep practical mastery of vLLM, Triton Inference Server, or TensorRT-LLM in production environments.
- Strong containerization experience (Docker, Kubernetes/K3s) and network routing (Nginx, Envoy).
Problem Solving, Execution Rigor & Curiosity
- Passionate about squeezing maximum FLOPS and memory bandwidth out of modern hardware.
- Zero tolerance for server downtime, resource leaks, or unmonitored infrastructure degradation.
- Builder who loves testing newly released inference optimizations the day they drop.
5-Day Live Technical Evaluation Milestone
5-Day Live Practical Milestone: Set up an optimized multi-GPU inference engine achieving a 2.5x throughput improvement and sub-15ms time-to-first-token under synthetic concurrent load (strictly 5 working days).
Institutional Hiring Protocol: Candidates who pass initial resume screening are invited to a live, practical evaluation milestone spanning strictly not more than 5 working days. Verifiable completion and code audit by your assigned senior engineering mentor is the sole prerequisite for official corporate offer letter issuance.
Compensation, Total Rewards & Advancement
- Generous annual package ₹12,00,000–₹21,00,000 with infrastructure uptime bonuses.
- Hands-on ownership of enterprise GPU hardware clusters.
- Remote-friendly team structure with flexible hours.
- Sponsorship for advanced Kubernetes and cloud architecture certifications.
Dedicated Inquiries Inbox for This Role
Have questions regarding architecture scope or wish to share private research repos directly? Messages sent to this address route straight to the engineering leads reviewing this opening.
cognitive-infrastructure-enginee-careers@cehpoint.co.in
Open Mail Client →