Benchling is building Intelligence Engineering & Enablement within its Security & IT organization to own internal AI tooling, cross-functional agentic AI applications, and the data foundations they depend on. As the founding engineer for this team, you'll architect and deliver enterprise-grade agentic AI systems that bridge departmental experimentation and production workloads.
About the Role
This is a hands-on AI systems engineering role focused on building reliable production systems around modern foundation models—not model research or ML training from scratch. You'll be a player-coach on a flat team, spending at least half your time writing code while leading technical direction, architecture, and delivery of Benchling's agentic AI portfolio. You'll partner closely with the AI Product Manager on prioritization, the Data, Analytics & Systems team on data foundations, and drive technical hiring and mentorship.
Responsibilities
Define foundational architecture for enterprise agentic AI—orchestration, agent frameworks, tool integrations (including MCP), memory and state management, evaluation, and observability. Make clear build vs. buy decisions with documented rationale.
Write production code to build and ship the early agentic system portfolio yourself. Stand up CI/CD, testing, evaluation, and deployment infrastructure, graduating prototypes into hardened production systems under a "you build it, you run it" model.
Design for enterprise security from day one: multi-tenant isolation, secrets management, audit logging, payload encryption, role-based access controls, and human-in-the-loop controls. Partner with Security Engineering on threat modeling for agentic architectures including prompt injection, tool misuse, and data exfiltration vectors.
Enable builders across the company by coaching power users and departmental teams on production patterns, developing criteria for prototype graduation, and building internal developer experience—templates, SDKs, sandboxes—for safe deployment.
Partner with Data, Analytics & Systems team on source-of-truth datasets and pipelines that agentic systems depend on. Engage department leaders on workflow transformation and leverage existing platform and infrastructure capabilities.
Set the bar for code quality, testing and evaluation, documentation, and on-call practices. Drive technical hiring through interview design and represent the team to senior candidates. Mentor engineers on the team and other AI builders across Benchling.
Requirements
7+ years of professional software engineering experience building production systems with strong systems design fundamentals.
Hands-on experience building production systems that integrate with LLMs and/or agentic patterns: orchestration, tool use, memory and state management, evaluation, and observability.
Demonstrated understanding of how to optimize workloads across deterministic and non-deterministic capabilities, striking the right architectural balance for specific solutions.
Production experience with at least two of: Python, TypeScript/Node.js, Go; comfort working across the stack.
Hands-on expertise with LLM APIs (OpenAI, Anthropic), agentic frameworks (LangChain, CrewAI), RAG over business content (Confluence, contracts, policies), vector databases (pgvector, Pinecone), workflow automation (n8n, Langflow), and LLM observability and evaluation tooling (LangSmith, Arize).
Track record of building platforms or product areas from zero to one and scaling them.
Experience operating in regulated or security-sensitive environments with solid grasp of enterprise security fundamentals—encryption, access controls, audit logging, secrets management.
Comfortable exercising technical leadership independent of positional authority. You set direction, raise the bar in design reviews, and grow other engineers through influence.
Product-first approach: you ship code quickly and care about real-world impact.
Strong communication skills with both technical and non-technical audiences. You translate department workflows into engineering plans and engineering tradeoffs into business language.
Interest in learning more about life science (prior knowledge not required).
Nice to Have
Background in enterprise SaaS, life sciences, or biotech.
Familiarity with LLM orchestration patterns and frameworks (LangGraph, MCP, agent SDKs from major model providers).
Experience with async orchestration (Temporal, Prefect, Airflow) applied to long-running or agentic workflows.
Familiarity with SOC 2, HIPAA, or GxP compliance as they apply to AI systems.
Experience building internal developer platforms or internal tools at scale.
Direct experience coaching or enabling non-engineers (analysts, ops staff, business power users) to build with AI tooling.
Based in or willing to relocate to San Francisco, CA (relocation assistance offered); remote US candidates also considered.
Benchling on Oh My Job
5 open positions right now, including 4 in California.