AI Engineer - Generative AI and Agents
Azumo · Remote · Argentina · Remote
Posted Sep 21, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Azumo builds and operates production AI systems for companies ranging from seed-stage startups to Meta. We are hiring an AI Engineer to own what those systems do once they are live: retrieval pipelines, tool-using agents, evaluation harnesses, and the guardrails that keep them dependable in front of real users. The role is fully remote across Latin America , aligned to your client's working day.
You will not be building demos. Azumo has shipped production AI since 2016, and the work here starts where the prototype ends, making a system reliable, measurable, and affordable enough to put in front of customers.
Where this role sits
Azumo's engineering organization is built around four lanes. The Data Engineer lane owns pipelines, storage, and the retrieval layer. The Data Scientist lane owns the question and the method. The Software Engineer lane owns AI-augmented product delivery. This role is the AI Engineer lane, and it owns production behavior.
One question places the boundary: when the output is wrong, whose problem is it? "The method was inappropriate" is a Data Scientist question. "The system did the wrong thing with an appropriate method" is yours.
Not quite your profile? Check our other openings:
If you build the pipelines and retrieval layer models depend on — Data Engineer If you decide what to measure and which method answers it — Data Scientist If you ship product software with agents in your toolchain — AI-Augmented Software Engineer If you've done all of the above and answered to the client directly — Forward Deployed Engineer
What you will build
Retrieval systems. Chunking and embedding pipelines, hybrid search, reranking, and evaluation of retrieval quality, built on pgvector, Pinecone, Qdrant, FAISS, or Azure AI Search.
Agentic workflows. Stateful multi-step execution, tool calling, MCP servers, structured output enforcement, context-window management, deterministic fallbacks, and human-in-the-loop gates for the decisions that need one.
Evaluation. Test sets that reflect the decision the system is actually making, model-as-judge scoring, regression tracking across prompt and model changes, and honest error analysis. If a change made the system better, you should be able to prove it.
Reliability and safety. Prompt-injection defense, output validation, guardrails, PII handling, and graceful degradation when a model or tool call fails.
Production operation. Containerized deployment on Azure or AWS, CI/CD, observability, and explicit latency, cost, and token budgets that you own rather than discover after the invoice.
Work inside the client's environment. Their repositories, their standups, sometimes their customer calls. Azumo is SOC 2 certified, client code stays in client repositories, and some engagements carry additional requirements such as HIPAA.
How we work
Our engineers build with AI every day. Claude Code, Codex, and similar tools are part of the standard toolchain here, not an experiment. We run an…