LLMOps Engineer
2026-08-04T11:22:10+00:00
ALX
https://cdn.greatkenyanjobs.com/jsjobsdata/data/employer/comp_3838/logo/ALX.png
https://www.alxafrica.com/
FULL_TIME
Nairobi
Nairobi
00100
Kenya
Telecommunications
Science & Engineering, Computer & IT, Engineering / Technical
2026-08-11T17:00:00+00:00
TELECOMMUTE
8
About ALX
At ALX, we’re unlocking the future we want to see. We’re catalyzing the transformation of Africa, by developing the next generation of bold, innovative, ethical and entrepreneurial leaders. We’re unlocking the potential of the world's largest workforce. The future is calling: be the answer.
Role Summary
Project A is ALX’s AI learning platform — a set of LLM products used by learners. Every one generates a stream of LLM data, and every one has hypotheses baked into it about what “working” means. The LLMOps Engineer owns the analyzer function: turning that stream into an honest answer about whether the products work. Take RAG as one example — documents must be stored accurately, fetched accurately, and fetched in the right mixture: three separate failure modes, each needing its own eval. Every product decomposes like that. This is a junior-to-mid role with a deliberate growth path: you start close to the technical lead’s designs and grow into full ownership of the function.
You will work in collaboration with Anthropic Engineers, a cross functional team of AI engineers, product managers and data scientists to design world class learning experiences.
Specific Responsibilities
Evaluation Suites
- Build and run eval suites per product, decomposed by failure mode, running on schedule and on every release — regression testing so nothing ships if it broke what worked.
- Keep evals cost-effective as the product line grows.
Reporting, Data & Collaboration
- Own the reporting loop — findings from evals and platform data in front of the team and stakeholders, including surfacing unintended or problematic model behaviour before learners do.
- Steward the core datasets the team depends on, including classified customer-support data.
- Partner with the AI Product Manager on instrumentation — they instrument the product, you build the evals over what is captured. This is a measurement role, not infrastructure — no model hosting or serving.
Skill Requirements - Essential
- Python & data: solid Python and a data inclination, comfortable shaping and analysing messy LLM-generated data.
- Decomposition: the ability to look at an AI product and decompose it into success and failure metrics.
- Eval landscape: familiarity with Langfuse, RAGAS, DSPy, or similar — depth in one, awareness of the rest. These tools are learnable; we hire the fundamentals underneath them.
Desirable (not required):
- experience keeping evals cheap at scale; dashboarding and reporting; classical statistics.
Essential Traits for Success
- You want to own a function, not execute tickets.
- You communicate well and like collaborating, you will support every builder on the team.
- You can point to any project, even a small one, where you measured an AI system honestly.
- Build and run eval suites per product, decomposed by failure mode, running on schedule and on every release — regression testing so nothing ships if it broke what worked.
- Keep evals cost-effective as the product line grows.
- Own the reporting loop — findings from evals and platform data in front of the team and stakeholders, including surfacing unintended or problematic model behaviour before learners do.
- Steward the core datasets the team depends on, including classified customer-support data.
- Partner with the AI Product Manager on instrumentation — they instrument the product, you build the evals over what is captured. This is a measurement role, not infrastructure — no model hosting or serving.
- Solid Python and a data inclination, comfortable shaping and analysing messy LLM-generated data.
- Ability to look at an AI product and decompose it into success and failure metrics.
- Familiarity with Langfuse, RAGAS, DSPy, or similar — depth in one, awareness of the rest.
- Experience keeping evals cheap at scale (desirable).
- Dashboarding and reporting (desirable).
- Classical statistics (desirable).
JOB-6a71cb621745d
Vacancy title:
LLMOps Engineer
[Type: FULL_TIME, Industry: Telecommunications, Category: Science & Engineering, Computer & IT, Engineering / Technical]
Jobs at:
ALX
Deadline of this Job:
Tuesday, August 11 2026
Duty Station:
This Job is Remote
Summary
Date Posted: Tuesday, August 4 2026, Base Salary: Not Disclosed
Similar Jobs in Kenya
Learn more about ALX
ALX jobs in Kenya
JOB DETAILS:
About ALX
At ALX, we’re unlocking the future we want to see. We’re catalyzing the transformation of Africa, by developing the next generation of bold, innovative, ethical and entrepreneurial leaders. We’re unlocking the potential of the world's largest workforce. The future is calling: be the answer.
Role Summary
Project A is ALX’s AI learning platform — a set of LLM products used by learners. Every one generates a stream of LLM data, and every one has hypotheses baked into it about what “working” means. The LLMOps Engineer owns the analyzer function: turning that stream into an honest answer about whether the products work. Take RAG as one example — documents must be stored accurately, fetched accurately, and fetched in the right mixture: three separate failure modes, each needing its own eval. Every product decomposes like that. This is a junior-to-mid role with a deliberate growth path: you start close to the technical lead’s designs and grow into full ownership of the function.
You will work in collaboration with Anthropic Engineers, a cross functional team of AI engineers, product managers and data scientists to design world class learning experiences.
Specific Responsibilities
Evaluation Suites
- Build and run eval suites per product, decomposed by failure mode, running on schedule and on every release — regression testing so nothing ships if it broke what worked.
- Keep evals cost-effective as the product line grows.
Reporting, Data & Collaboration
- Own the reporting loop — findings from evals and platform data in front of the team and stakeholders, including surfacing unintended or problematic model behaviour before learners do.
- Steward the core datasets the team depends on, including classified customer-support data.
- Partner with the AI Product Manager on instrumentation — they instrument the product, you build the evals over what is captured. This is a measurement role, not infrastructure — no model hosting or serving.
Skill Requirements - Essential
- Python & data: solid Python and a data inclination, comfortable shaping and analysing messy LLM-generated data.
- Decomposition: the ability to look at an AI product and decompose it into success and failure metrics.
- Eval landscape: familiarity with Langfuse, RAGAS, DSPy, or similar — depth in one, awareness of the rest. These tools are learnable; we hire the fundamentals underneath them.
Desirable (not required):
- experience keeping evals cheap at scale; dashboarding and reporting; classical statistics.
Essential Traits for Success
- You want to own a function, not execute tickets.
- You communicate well and like collaborating, you will support every builder on the team.
- You can point to any project, even a small one, where you measured an AI system honestly.
Work Hours: 8
Experience in Months: 12
Level of Education: bachelor degree
Job application procedure
Click Here to Apply Now
All Jobs | QUICK ALERT SUBSCRIPTION