FULL-TIME CORPORATE APPOINTMENT Infrastructure & Cloud Systems 100% REMOTE (INDIA)

Cognitive Infrastructure Engineer

Optimize high-throughput LLM serving infrastructure, tensor parallel clusters, vLLM / TensorRT-LLM runtimes, dynamic batching, and hardware-accelerated model routing.

Compensation ₹12,00,000 – ₹21,00,000 / year
Location & Work Model Remote (India)
Experience Requirement 2–5 Years
Hiring Benchmark 5-Day Code Audit
SECTION 01

Role Overview & Operational Scope

AI models are only as good as the infrastructure serving them. You will ensure Cehpoint's proprietary models and agent clusters run with rock-solid reliability, sub-15ms time-to-first-token, and peak GPU utilization across cloud and on-premise hardware deployments.

vLLM / TensorRT-LLMCUDA & GPU ClusteringDocker / KubernetesLinux Kernel TuningPrometheus & GrafanaPython / C++Model Quantization (AWQ/GGUF)
SECTION 02

Key Responsibilities & Production Deliverables

  • Deploy, scale, and optimize inference engines (vLLM, TensorRT-LLM, TGI) across multi-GPU NVIDIA clusters.
  • Implement PagedAttention, continuous batching, and speculative decoding to double token generation throughput.
  • Configure automated model caching, weights distribution, and failover across regional edge locations.
  • Design comprehensive Prometheus/Grafana monitoring dashboards tracking KV cache usage, queue depths, and token latency.
  • Maintain rock-solid security hardening across cloud VPS instances, container runtimes, and VPC networks.
SECTION 03

Mandatory Foundational Knowledge

  • GPU microarchitecture, CUDA execution hierarchy, high-bandwidth memory (HBM), and NVLink interconnects.
  • KV-cache mechanics, tensor parallelism, pipeline parallelism, and weight quantization techniques.
  • Linux kernel networking, container isolation, and enterprise security compliance.
SECTION 04

Mandatory Practical Skills & Architecture

  • Expertise in Linux system administration, bash automation, and infrastructure-as-code.
  • Deep practical mastery of vLLM, Triton Inference Server, or TensorRT-LLM in production environments.
  • Strong containerization experience (Docker, Kubernetes/K3s) and network routing (Nginx, Envoy).
SECTION 05

Problem Solving, Execution Rigor & Curiosity

  • Passionate about squeezing maximum FLOPS and memory bandwidth out of modern hardware.
  • Zero tolerance for server downtime, resource leaks, or unmonitored infrastructure degradation.
  • Builder who loves testing newly released inference optimizations the day they drop.
SECTION 06 · PRACTICAL EVALUATION BENCHMARK

5-Day Live Technical Evaluation Milestone

5-Day Live Practical Milestone: Set up an optimized multi-GPU inference engine achieving a 2.5x throughput improvement and sub-15ms time-to-first-token under synthetic concurrent load (strictly 5 working days).

Institutional Hiring Protocol: Candidates who pass initial resume screening are invited to a live, practical evaluation milestone spanning strictly not more than 5 working days. Verifiable completion and code audit by your assigned senior engineering mentor is the sole prerequisite for official corporate offer letter issuance.

SECTION 07

Compensation, Total Rewards & Advancement

  • Generous annual package ₹12,00,000–₹21,00,000 with infrastructure uptime bonuses.
  • Hands-on ownership of enterprise GPU hardware clusters.
  • Remote-friendly team structure with flexible hours.
  • Sponsorship for advanced Kubernetes and cloud architecture certifications.
SECTION 08 · DIRECT INQUIRIES & CATCH-ALL

Dedicated Inquiries Inbox for This Role

Have questions regarding architecture scope or wish to share private research repos directly? Messages sent to this address route straight to the engineering leads reviewing this opening.

cognitive-infrastructure-enginee-careers@cehpoint.co.in Open Mail Client →

Related Engineering Appointments

CORPORATE

AI Personality System Architect

Remote (India) · ₹14,00,000 – ₹24,00,000 / year + Equity Options
View Specification →
CORPORATE

AI Instinct Researcher

Remote (India) · ₹12,00,000 – ₹22,00,000 / year + Research Grant
View Specification →
CORPORATE

AI Memory & Knowledge Graph Engineer

Remote (India) · ₹11,00,000 – ₹19,00,000 / year
View Specification →