About us
Instant Systems builds and scales technology ventures out of its development labs in India. One of them, InstantMarkets, is an open business search engine for procurement - it aggregates government and private-sector bids, RFPs, RFQs and contract awards from thousands of sources worldwide and makes them searchable in one place. Across the portfolio, our products share a common problem: enormous volumes of messy, unstructured, real-world documents that need to be understood rather than merely stored.
The role
InstantMarkets and other Instant Systems products sit on large, messy corpora: bids, RFPs, RFQs, awards, and enterprise documents that have to be retrieved, reasoned over, and kept trustworthy. We are hiring Applied AI Engineers to design and ship complete systems — ingestion, retrieval, agent/tooling, serving, evaluation, and cost control — not thin LLM wrappers or one-off prompt improvements.
You will use coding agents (Claude Code, Codex, Cursor) as the default way of working, and you will be accountable for what they produce: tests, security, evals, token/cost, and operability. Some work is product-core; some is embedded with a customer or portfolio team until the system is in production and handed over.
The interview includes a practical assignment: download data and build an application with APIs and a CLI on top — the same shape of work you will do here.
How we assess (take-home assignment)
Shortlisted candidates complete a timed take-home. You will download a real dataset, then build a working application on top of it: HTTP APIs, a CLI, and a small app or service that uses both. This is not a prompt-engineering exercise and not a slide deck.
We expect you to stand up an end-to-end slice: ingest the files, expose query/command interfaces, and make the system usable from the terminal and over the API. Use coding agents if you want (Claude Code, Codex, Cursor). You still own design, correctness, and what you submit.
What we look at:
- Can you go from raw download → structured data → running app without hand-waving?
- Are the APIs and CLI real interfaces (args, flags, help, error handling, not a single happy-path script)?
- Do you make sensible choices for retrieval, tools, or automation only where they help, and can you explain cost/quality tradeoffs?
- Is there a README that another engineer can follow: how to download, run, call the API, and run the CLI?
We will share the brief, data access instructions, timebox, and deliverable list with candidates who reach this stage. Senior candidates should expect a slightly harder brief (clearer bar on architecture, eval, and operability).
What you'll do
- Design, build, and deploy AI-native software solutions for strategic customers.
- Lead technical delivery from architecture and development through deployment and handover.
- Work across backend services, APIs, data pipelines, integrations, and AI components.
- Embed with customer teams to understand their technical environment, requirements, and success criteria.
- Use coding agents such as Claude Code, OpenAI Codex, Cursor, and GitHub Copilot to accelerate development.
- Design and build multi-agent systems, orchestration workflows, and reusable agent components.
- Integrate LLM APIs and build RAG pipelines using embeddings, vector databases, and retrieval strategies.
- Build reliable AI systems with testing, guardrails, fallback logic, audit logging, and human-in-the-loop controls.
- Review and validate AI-generated code to ensure security, reliability, and enterprise-grade quality.
- Identify reusable technical patterns and customer pain points to help improve the product and platform.
What we're looking for
- Own a production path: document ingest → chunk/index → hybrid retrieval → reasoning/tools → API or UI → monitoring.
- Build retrieval that is measured (recall, faithfulness, latency), not “the answer looked fine.”
- Design agent/tool workflows only where a pipeline is not enough: tool calling, MCP connectors, human-in-the-loop, durable state.
- Put systems on a real host (PaaS or containers): services, workers, Postgres, and a search/vector layer.
- Instrument quality and spend: traces, eval sets, guardrails, fallbacks, tokens per task / cost per request.
- Direct coding agents with skills, hooks, and review; do not paste unreviewed agent output into production.
- Sit with customers or internal product teams, turn ambiguity into an architecture and a shipped increment.
- Extract reusable platform pieces (connectors, eval harness, index patterns) back into Instant Systems.
Responsibilities
- Design, build, and deploy AI-native software solutions for strategic customers.
- Lead end-to-end technical delivery from architecture and development through deployment and handover.
- Work directly with customer teams to understand their technical environment, requirements, and success criteria.
- Design multi-agent systems, orchestration workflows, and reusable agent components.
- Integrate LLM APIs and build production-ready RAG pipelines using embeddings and vector databases.
- Implement testing, guardrails, fallback logic, security, audit logging, and human-in-the-loop controls.
Must Have
Production Python or TypeScript
- At least one end-to-end AI system they can walk: data → retrieve or tools → serve → evaluate.
- RAG or search over messy documents: embeddings, hybrid search, a real index (pgvector, Qdrant, Pinecone, Weaviate, OpenSearch, Turbopuffer/LanceDB, or equivalent).
- One orchestration or tool layer they actually used: LangGraph, Pydantic AI, Agents SDK, MCP servers/tools — not “built a chatbot.”
- Production delivery: APIs, a database, deploy (Railway/Render/Fly/Cloud Run/Vercel+backend, or K8s).
- Failure modes: hallucinations, stale retrieval, tool errors, fallback, audit trail.
- Daily use of an AI coding agent and evidence they review/govern it.
- Something inspectable: GitHub, internal design doc, metrics (latency, cost, eval score).
- Chose when not to use an agent (pipeline vs graph vs single LLM call).
- Owned eval and cost in production (LangSmith/Langfuse/Phoenix/RAGAS or equivalent).
- Designed tenancy, deletion, or memory-vs-RAG if the system personalizes.
- Turned one customer build into a reusable pattern.
Nice to have
- Agent memory layer (Mem0, Zep/Graphiti, Letta, LangMem) and can explain memory ≠ RAG.
- Written SKILL.md / hooks / MCP servers, not only installed a pack (ECC, Superpowers, spec-kit).
- Devcontainers, Codespaces, or an internal harness for team DevEx.
- Experience with CI/CD, Kubernetes, Terraform, or cloud infrastructure.
- Background in consulting, solutions engineering, or startup environments.
- Contributions to open-source AI/agent tooling or published work on AI system design.
- Fine-tuning — only if eval is already solid.
What's great in the job?
- Work on cutting-edge AI-powered products and enterprise applications.
- Build AI-native solutions using Generative AI, LLMs, autonomous agents, and intelligent automation.
- Work directly with customers on complex technical challenges and real-world use cases.
- Collaborate with talented engineers and global product teams.
- Grow through mentorship, technical leadership, and continuous learning.
- Work in a flexible environment that promotes work-life balance.
- Be part of an inclusive culture where diversity, innovation, and collaboration are valued.
- See your engineering work create meaningful business impact.
What We Offer
Each employee has a chance to see the impact of his work. You can make a real contribution to the success of the company.
AI Innovation
Work on cutting-edge AI-powered products and enterprise applications.
Flexible Environment
Work in a flexible environment that promotes work-life balance.
Global Collaboration
Collaborate with talented engineers and global product teams.
iThrive Career Growth
Build your skills, receive continuous feedback, and grow through a structured career development path.