Loading jobs…
Loading jobs…
Toloka — European district, Brussels Capital
About Toloka At Toloka AI we create data that powers leading GenAI models and innovations. We work with frontier labs, big tech, renowned AI startups, enterprises and non-profit research organizations worldwide. We use a combination of Experts + Crowd + Tech Platform to teach AI models to reason and evaluate their efficacy and safety.
We have experts in more than 50 different domains—from doctors and lawyers to physicists and engineers—and boast one of the most diverse global crowds, representing over 100 countries and speaking 40+ languages . We are a well-funded startup with an enviable portfolio of clients including Anthropic , Amazon , Microsoft , Poolside , Recraft , and Shopify . Recently, we secured strategic investment led by Bezos Expeditions and Nebius Group with participation from Mikhail Parakhin , CTO of Shopify and board advisor to leading GenAI companies, who now serves as our Chairman of the Board.
Our remote-first team is globally distributed around the world: USA, UK, the Netherlands, Serbia , and more. About the Position Led applied ML initiatives within the Delivery division, integrating LLM and AI-agent technologies into active client engagements with a focus on Anti-Fraud. Optimize output quality and automate manual workflows to scale operational capacity and improve business margins.
Core Outcomes You’ll Drive Automation lift: Achieve double-digit percentage improvements per account by replacing manual steps with robust LLM/agent pipelines. Quality & reliability: Raise client acceptance rates and reduce rework through automated LLM checks and evaluation harnesses. Throughput & margin: Shorten cycle times and expand task capacity by productizing reusable components across projects.
Safety & governance: Enforce strict guardrails and solution auditability through systematic red-teaming.
What You’Ll Do
Solution architecture: Design and deploy agentic workflows (tool use, planning, retrieval, critique loops) for data generation, evaluation, anti-fraud, and safety use cases. Automated evaluation: Build judge models, auto-grading systems, and self-verification pipelines to codify acceptance criteria into repeatable checks. Productization: Create reusable reference architectures, templates, and components to scale delivery across multiple client accounts.
AI-driven scaling: Integrate LLMs into contributor workflows (pre-label → verify → escalate) to eliminate operational toil and maximize expert leverage. Observability: Implement production measurement frameworks (task KPIs, drift), online A/B tests, and cost/latency dashboards.
Requirements
into pragmatic ML designs. Technical stewardship: Set engineering standards for prompts, agents, data pipelines, and CI/CD via active code and design reviews.
What We'Re Looking For
Production experience: Over 5–8+ years of expertise in applied ML and LLM development, with a documented history of launching agentic workflows into production. Domain expertise: Strong background in building anti-fraud platforms, trust & safety systems, or real-time anomaly detection frameworks. Delivery focus: A pragmatic approach, capable of transforming vague project needs into executable solutions defined by strict timelines and measurable KPIs.
AI quality assurance: Extensive experience in building automated evaluation pipelines and grading systems to maintain consistency at scale. Automation mindset: A natural drive for automation, with a track record of identifying and optimizing manual operations to enhance throughput and margins. System architecture: Skill in pragmatic engineering, prioritizing simple, observable, and efficient designs that align with latency and cost goals.
Hands-on engineering: A technical profile featuring advanced Python proficiency and the ability to tune agents and prompts within live systems personally.