AI Behavior & Alignment Architect
Formulate formal behavioral boundaries, automated red-teaming harnesses, constitutional AI principles, and steerability systems for mission-critical client deployments.
Role Overview & Operational Scope
As autonomous agents gain access to enterprise databases, shell environments, and transactional systems, behavioral failure is catastrophic. You will engineer mathematical boundaries and runtime monitors that guarantee agent actions remain compliant with corporate governance and legal mandates.
Key Responsibilities & Production Deliverables
- Author formal Constitutional AI principles and runtime guardrails for autonomous tool-execution loops.
- Design automated adversarial attack simulations (jailbreaks, prompt injections, privilege escalations, context contamination).
- Build real-time behavioral interceptors that inspect tool arguments before OS/API execution occurs.
- Publish formal System Cards and Risk Assurance Dossiers for enterprise client boards and regulatory compliance audits.
- Establish automated regression testing suites that prevent behavioral degradation during model upgrades.
Mandatory Foundational Knowledge
- In-depth knowledge of LLM vulnerability taxonomies: OWASP Top 10 for LLMs, NIST AI RMF, and MITRE ATLAS.
- Theoretical understanding of value alignment, reward hacking, specification gaming, and goal misgeneralization.
- Familiarity with regulatory frameworks including India's DPDP Act, EU AI Act, and enterprise ISO standards.
Mandatory Practical Skills & Architecture
- Proficiency in Python for building automated security test harnesses and fuzzing pipelines.
- Hands-on experience deploying guardrail toolkits (NVIDIA NeMo Guardrails, Llama Guard, custom AST parsers).
- Strong technical writing skills for authoring institutional safety reports.
Problem Solving, Execution Rigor & Curiosity
- Relentless adversarial mindset that anticipates edge cases before they manifest in production.
- Commitment to practical, high-throughput security that does not needlessly strangle agent utility.
- Continuous research into newly published alignment failures and prompt injection vectors.
5-Day Live Technical Evaluation Milestone
5-Day Live Practical Milestone: Construct an automated adversarial fuzzing suite that tests autonomous agents against jailbreaks, unauthorized system execution, and data exfiltration vectors, generating automated compliance reports (strictly 5 working days).
Institutional Hiring Protocol: Candidates who pass initial resume screening are invited to a live, practical evaluation milestone spanning strictly not more than 5 working days. Verifiable completion and code audit by your assigned senior engineering mentor is the sole prerequisite for official corporate offer letter issuance.
Compensation, Total Rewards & Advancement
- Annual compensation ₹13,00,000–₹22,00,000 + performance bonus.
- High visibility role presenting directly to enterprise CISOs and engineering executives.
- Fully remote setup with dedicated research compute allocation.
- Generous annual training budget for advanced security and AI certifications.
Dedicated Inquiries Inbox for This Role
Have questions regarding architecture scope or wish to share private research repos directly? Messages sent to this address route straight to the engineering leads reviewing this opening.
ai-behavior-alignment-architect-careers@cehpoint.co.in
Open Mail Client →