NinjaTech AI is a generative AI startup (B2C and B2B) with headquarters in Silicon Valley and offices in Sydney and Vancouver. Backed by Alexa Fund and Samsung Ventures, we are building the autonomous AI agents that let anyone get real work done with generative AI.
Our flagship product, SuperNinja, is an advanced agentic AI platform with full OS capabilities: website creation, end-to-end coding, advanced data analysis, and more. The three roles below build the agent harness itself - the runtime, integrations, and reliability layer that every SuperNinja agent runs on.
Position Overview
We are looking for a Core Runtime engineer to own the agent harness that turns our LLMs into reliable, autonomous software. The harness is the run loop, orchestrator, and monitor that keep an agent working safely across thousands of tool calls and hours of wall-clock time. You will work at the intersection of systems engineering and applied AI, building the substrate that every SuperNinja agent runs on.
Key Challenges & Opportunities
Design and evolve the agent run loop: context management, turn compaction, and recovery from failed or partial tool calls
Build the orchestrator that launches, supervises, and hands off long-running background agents without losing state
Implement the monitoring layer that detects stuck, looping, or runaway agents and intervenes safely
Own the provider abstraction across model backends (Claude, Codex, and others) so the harness stays model-agnostic
Make the harness observable end to end with structured telemetry, so failures are debuggable in production
Keep the loop fast and cheap at scale while it powers millions of daily user interactions
Experience & Education Requirements
Master's or Bachelor's in Computer Science or equivalent practical experience
4+ years building production backend or systems software in Python
Strong grasp of concurrency, process supervision, and fault-tolerant design
Experience with LLM APIs and agentic patterns (tool calling, function calling, streaming)
Track record of shipping and operating reliable services at scale
About the Role
You will report to the engineering lead for agent infrastructure and will have ownership in these areas:
Research, design, and build the core run loop and orchestration that power autonomous agents
Define how agents manage context, recover from errors, and hand off work cleanly
Build rapid prototypes and proof of concepts to turn harness ideas into shipped product capabilities
Stay up-to-date with agent architectures from research and industry, and fold the best ideas into our runtime
Design self-healing mechanisms so agents degrade gracefully instead of failing hard
Collaborate with cross-functional teams including applied science, product, and integrations
What Sets You Apart
You have built an agent framework, workflow engine, or long-running job system before
You care about reliability the way SREs do, and you instrument everything
You can reason about non-determinism in LLM outputs and design around it
Experience with sandboxing, containerization, or secure code execution
You move fast on prototypes but know when to harden them for scale
Why Join Us
This is a chance to build the foundation that every one of our agents runs on. The harness is where autonomy is won or lost, and you will own it end to end alongside a team that values engineering rigor and scientific curiosity.