Oh My JobFind Jobs
Company
  • About Us
  • Blog
  • Contact
Tools
  • Paycheck Calculator
Legal
  • Terms of Service
  • Privacy Policy
  • California Privacy Rights
For Employers
Post a Job

Senior Applied AI Engineer

Level
Level
Austin
May 26, 2026
Salary not listed
FullTime

Job Description

Senior Applied AI Engineer role at Level, a learning technology company based in Austin. You'll design and ship production agentic systems, automate high-leverage workflows with AI, and build the evaluation frameworks that determine whether features launch. This is hands-on work: you own prompts, tools, retrieval, guardrails, observability, and the complete handoff to production teams.

About the Role

Level builds interactive learning experiences for students, teachers, and parents. As a Senior Applied AI Engineer, you'll work onsite in Austin (5 days per week; relocation offered) to architect and deploy AI agents and agentic workflows that solve defined problems end-to-end. You'll partner with subject matter experts to replace operational toil with trustworthy AI systems, own the measurement and evaluation frameworks that gate launches, and ensure AI features meet compliance and safety standards before they reach end users.

Responsibilities

  • Design and build production agentic systems and workflows that solve concrete problems. Own prompts, tool integration, retrieval pipelines, guardrails, observability, cost and latency budgets, and rollout strategy.
  • Identify and automate high-leverage operational workflows (review pipelines, content QA, labeling, support triage, internal copilots) by partnering with SMEs to define success criteria, validate outputs, and iteratively replace manual work with AI systems they trust.
  • Build and maintain the evaluation infrastructure: offline evaluations, LLM-as-judge with calibration, regression test suites, and online production metrics. Create and version golden datasets covering intents, difficulty levels, edge cases, and adversarial inputs; continuously enrich them with production failures and SME-validated labels.
  • Design AI features as regulated products. Implement compliance controls relevant to K-12 education (COPPA, FERPA), run safety and bias evaluations before launch and continuously post-launch, and build human-in-the-loop and content-filtering safeguards.
  • Deliver production-ready handoffs: documentation, runbooks, evaluation harnesses, and monitoring dashboards. Embed with receiving teams (4–12 weeks typical) until they ship changes independently.
  • Create internal libraries, patterns, and playbooks that enable other engineering teams to build and ship AI features without direct involvement.

Requirements

  • 7+ years of professional engineering experience with at least 1+ years of hands-on production work building agentic systems using a modern framework (LangGraph, Anthropic SDK, OpenAI Agents SDK, Pydantic-AI, Mastra, LlamaIndex, CrewAI, or equivalent homegrown stack).
  • Strong Python proficiency plus fluency in one typed language for production services (TypeScript, Go, or similar). Cloud platform experience (AWS or GCP) and containerized deployment.
  • Senior or staff-level software engineering foundation with multiple years of production environment experience and a track record of leading systems to launch.
  • Multiple shipped LLM-powered features in production with concrete stories about failures, fixes, and lessons learned.
  • Practical knowledge of agentic patterns: ReAct, tool use with structured schemas, prompt chaining, routing, orchestrator-worker, evaluator-optimizer/reflection, and human-in-the-loop. Judgment to choose deterministic workflows over autonomous agents when appropriate.
  • Production retrieval system experience: chunking, embeddings, hybrid search, re-ranking, metadata filtering, and their failure modes. Working knowledge of grounding techniques (citation/quote extraction, faithfulness evals, refusal evals, consistency checks).
  • Strong prompt engineering practice: zero-shot, few-shot, and many-shot patterns; example selection and ordering; in-context learning and chain-of-thought.
  • Hands-on experience with structured output and validation in production (provider-native structured outputs, Instructor, Pydantic-AI, Outlines, or equivalent).
  • Disciplined evaluation practice. You make launch decisions based on objective measurement, not subjective review.
  • Strong written and verbal communication. You can explain architectural trade-offs clearly to executives and junior engineers.
  • Comfort using AI coding tools heavily in implementation while owning problem framing, design decisions, and verification. Success is measured by working systems shipped, not lines of code.

Nice to Have

  • Advanced retrieval experience: GraphRAG, agentic retrieval, evaluation-driven retrieval tuning, or hybrid retrieval at scale.
  • Direct or transferable experience with safety, privacy, and policy constraints in user-facing AI systems. K-12 or other regulated-domain background is a strong plus.
  • Production experience with prompt-optimization frameworks (DSPy, TEXTGRAD, AdalFlow).
  • Public repository, package, gist, or technical write-up demonstrating meaningful AI work, or a representative project you can describe in detail under confidentiality constraints.
  • Open-source contributions to AI tooling: frameworks, agents, evaluation tools, or MCP servers.

Interview Process

One interview round is an AI-assisted coding session. Bring your own setup (IDE, AI coding assistant, agentic tools—whatever you use daily) and solve a realistic problem live with AI in the loop. We evaluate how you collaborate with AI: prompting, validating output, catching bad suggestions, overriding when needed, and delivering production-quality code. This is not a cleanroom algorithm test.

Level on Oh My Job

4 open positions right now, including 2 in Texas.

Apply now
Share: