Founder / AI automation engineer

Kartik BK Mishra

I build and evaluate AI agents, voice workflows and automation systems with a strong bias toward testable behavior, human control and practical handover.

Kartik BK Mishra

The operator behind AuraMind

AI engineering starts with the workflow, not the tool list.

I founded AuraMind to work on the difficult middle between an impressive AI demo and a system that a team can actually operate. That middle includes the source data, prompts, retrieval, tool permissions, failure states, approval points, logs, handoffs and ownership questions that determine whether an agent is useful after the first presentation.

My work spans AI automation engineering, model evaluation, rubric design, RAG workflows, voice agents, agent red teaming, Python-based tooling and public-safe product prototypes. I use tools such as Codex, Claude Code and current model platforms as engineering environments, not as the product story by themselves. The goal is to shape a reliable process on top of them: define the decision, make the evidence visible and keep accountable people in control of impactful actions.

This process orientation also comes from my academic background. I hold an M.Sc. in Supply Chain Management alongside a computer science background. Supply-chain thinking is useful in AI systems because both are networks of dependencies: one weak handoff, unclear owner or unavailable input can undermine the entire outcome. AuraMind therefore treats agent design as an operating-system problem as much as a model problem.

Current focus

Reliable agents and inspectable automation.

I am currently focused on four connected areas. First, agent reliability: scenario design, trace inspection, retrieval quality and regression-oriented evaluation. Second, voice workflows: multilingual intake, consent, structured request capture and human handoff. Third, automation engineering: connecting business tools without hiding provider dependencies or ownership. Fourth, AI training and evaluation work: rubric-based review, red-team thinking and quality-control systems that help teams understand model behavior.

The public projects linked from AuraMind are evidence of this direction. The Agentic Eval Ops Kit explores reusable scorecards. The Enterprise AI Evaluation Framework organizes workflow, retrieval, safety and audit-readiness checks. The Realtime Voice Agent QA Console examines the events around a call rather than treating a fluent voice as sufficient proof. Zara is the public voice reference implementation. TrendAtlas remains a research prototype until its data and user evidence justify stronger positioning.

These artifacts are intentionally described as engineering proof, not customer outcomes. I do not use repository counts, model-provider names or synthetic fixtures as substitutes for permissioned client evidence.

How I work

Map, build, break, hand over.

I start by defining one outcome and the action that cannot go wrong. Then I map the inputs, systems, permissions and accountable owner. The first build is deliberately narrow so the team can observe the real workflow rather than debate an abstract architecture. Testing follows quickly: common cases, ambiguous requests, missing tools, weak evidence, unsafe instructions and recovery paths.

The handover should make the system less mysterious. It states what source is included, what third-party services remain dependencies, which accounts and subscriptions the owner controls, how to rerun tests and when the agent must stop for a person. If the evidence is incomplete, the public wording stays conservative.

Public links

Inspect the current work directly.