We're looking for a Senior Agentic AI Engineer to design, build, evaluate, and continuously improve sophisticated AI systems for a fast-moving enterprise AI platform. You will work across multi-agent architectures, context graphs, memory systems, model routing, and automated evaluation frameworks to build production systems that support complex, real-world business decisions.
This is not an AI API integration or prompt-engineering role. We are looking for someone who thinks deeply about how models, agents, tools, context, memory, and evaluation systems work together and can turn that understanding into reliable production systems. The ideal engineer combines strong software engineering fundamentals with deep curiosity, experimentation, and hands-on experience building modern agentic AI systems. If that describes you, this role is a strong fit. Why You'll Want to Join
You will be paid in USD (bi-monthly: every 15th and 30th)
Paid Time Off in accordance with company policy
Observance of Holidays per company guidelines
100% remote setup so you can work wherever you're most productive
This role requires availability during US business hours
High ownership over AI architecture and technical direction at an early, high-leverage stage
Work directly with enterprise customers and see the real impact of what you build
What You'll Work On
Agentic AI Architecture
Design and build production-grade agentic and multi-agent systems
Architect specialized agents with clearly defined roles, tools, context, permissions, and decision boundaries
Design orchestration and routing strategies across agents, tools, models, and workflows
Build systems that intelligently use multiple LLMs and model providers depending on the task
Determine which models should power specific agents based on quality, reasoning capability, reliability, latency, and cost
Design failure handling, fallbacks, guardrails, and human-in-the-loop mechanisms
Agent and Context Graph Interactions
Design and manage how agents interact with context graphs, knowledge graphs, memory systems, and other context sources
Determine what information and which portions of a context graph individual agents should be able to access
Design context retrieval, filtering, ranking, and permission strategies based on each agent's role and task
Build and evaluate interactions between agents and nested or interconnected context graphs
Determine how much context an agent needs without unnecessarily increasing tokens, latency, or noise
Design systems that dynamically provide agents with the most relevant context for a specific task
Evaluate how changes to context availability affect agent accuracy, reasoning, reliability, and performance
AI Evaluation and Reliability
Design sophisticated evaluation frameworks for agentic AI systems
Build evals across model, agent, and context graph interactions
Systematically evaluate which models perform best for specific agents, tasks, and workflows
Evaluate which context, memory, or graph information should be available to different agents
Implement adversarial testing, model-to-model critique, LLM-as-judge, regression testing, and systematic evaluation
Identify failure modes and use evaluation results to continuously improve system architecture
Optimize AI systems across quality, accuracy, reliability, latency, token usage, and cost
Prefer measurable experimentation and evaluation over assumptions when making architectural decisions
Context, Memory, and Knowledge Systems
Design context and memory architectures for autonomous and semi-autonomous agents
Build with RAG, embeddings, retrieval systems, vector databases, memory systems, context graphs, and knowledge graphs
Design how information flows between context systems and individual agents
Determine how context should be retrieved, structured, filtered, and updated
Build context systems that support different agents, tasks, and levels of access
Balance context