
Hrishikesh Jadhav
AI Engineer
AI engineer in Germany with 4+ years shipping production AI and backend systems. I build LLM extraction pipelines, RAG and MCP systems, and the evaluation gates that decide what ships.
I'm Hrishi, and at Toku I built the LLM extraction pipeline, RAG assistant and MCP server that run in daily production inside a global payroll platform. Before that I was a data scientist at GfK / NIQ in Nuremberg, building LLM classification, semantic retrieval and prediction models and running them in production with Python, AWS, Docker and GitLab CI/CD.
Work
Toku help centre
Customer-facing help centre and AI assistant for payroll, benefits, policy and country guides.
Two separate retrieval collections, so immigration questions never resolve against payroll content. How it works
- Problem
- Give customers one place for payroll, benefits, policy and country-guide answers, with an assistant that answers from that content.
- Architecture
- A Next.js help centre built on Toku's Notion knowledge base. Its assistant retrieves from two separate collections, general help and visa support, with fallback across model providers.
- Hard decision
- Retrieval is split into two collections so that immigration and work-authorisation questions never resolve against general payroll content.
- Evidence
- Helped clients get answers from Toku's Notion knowledge base.
Toku status page
Public status page for Toku's payroll platform.
Probing the edge and the origin separately tells a CDN failure apart from an application failure, and code fixes go straight to an agent. How it works
- Problem
- Catch failures on the payroll platform, hand the ones that need a code fix to an agent automatically, and notify the team on Slack promptly.
- Architecture
- Monitored components are probed from the Cloudflare edge on a short fixed interval. Incidents are detected and resolved automatically, dependencies roll up into one view, and status is published as a JSON API and an RSS feed. When an incident opens, the team is notified on Slack, and failures that need a code fix are picked up by a coding agent.
- Hard decision
- The edge and the origin are probed separately, so a failure is localised to the CDN layer or the application before anyone looks at it.
Experience
- 10/2025–09/2026
AI Engineer, Toku
Global employer-of-record and payroll infrastructure platform operating across 46 countries.
- Shipped AI tooling into daily production across product, engineering and sales teams: an MCP server exposing internal systems to LLM clients, a RAG assistant (FastAPI, LangChain, pgvector), and n8n automations for onboarding, payslip processing and support triage. Owned deployment, documentation, onboarding and hands-on user support.
- Cut payroll data intake from hours of manual work per cycle to minutes across production records by building a self-hosted LLM extraction pipeline (Python, FastAPI, PostgreSQL, DigitalOcean) that masks personal data before inference.Self-hosted so that no employee data leaves controlled infrastructure.
- Kept field-level extraction accuracy above the agreed threshold across every extracted field, checked against a large golden evaluation set, by building the evaluation harness and a CI regression gate that blocks any deploy below it.So a change that lowers extraction accuracy cannot reach payroll data.
- Moved security review to pull-request time across thousands of pull requests by integrating automated LLM-assisted security review into GitHub CI for Toku's financial platform.
- Built status.toku.com, Toku's public status platform: production components probed from the Cloudflare edge on a short fixed interval, automated incident detection and resolution, dependency rollups, edge-vs-origin failure isolation, and a JSON API and RSS feed. New incidents notify the team on Slack, and failures that need a code fix are picked up by a coding agent.Probing the edge and the origin separately localises a failure to the CDN or the application before anyone looks at it.
- Built Toku's customer-facing help centre and AI assistant (toku.com/help, Next.js) on Toku's Notion knowledge base, covering payroll, benefits, policy and country-guide articles across dozens of countries, with general-help and visa-support content in separate retrieval corpora and multi-model fallback across providers.Separate corpora so immigration and work-authorisation questions never resolve against general payroll content.
- Brought LLM inference costs down to a fraction of the estimated cost of equivalent third-party API usage by deploying open-weight models on self-hosted GPU infrastructure.
Personal data is masked before inference, and a golden-set gate decides what deploys. - 08/2022–09/2025
Data Scientist, GfK - An NIQ Company
Working Student, Data Science, 08/2022–07/2024
- Reduced manual effort in product-taxonomy classification by around 30% by building an LLM classification service on Amazon Bedrock using prompt engineering and RegEx-based post-processing.
- Reduced catalogue lookup time by around 40% versus keyword search by building vector-embedding semantic retrieval over internal product catalogues.
- Improved simultaneous-viewer prediction accuracy by around 20% over the incumbent model by training and deploying a CatBoost model using 30+ engineered features from sociodemographics, temporal patterns and programme metadata.
- Kept production scoring pipelines running 24/7 without manual intervention by orchestrating S3 ingestion, feature generation, model scoring and automated integration tests in GitLab CI/CD, containerised with Docker on Linux.
- Built a TV-audience data-fusion pipeline using K-Nearest Neighbors and a genetic algorithm to match and clone panel households against large-scale return-path data.
- 04/2020–08/2020
Software Developer (Internship), Sapio Analytics Pvt. Ltd.
- Built COVID-19 decision-support models in Python (SEIRD, scikit-learn, SciPy) and forecasting dashboards (Plotly, AWS) used to inform Government of India lockdown and testing-strategy decisions, with a 4.8% RMSE reduction against the prior baseline.
Research
- 2025
- 2021
Awards
- 2024
NIQ/GfK HACKFEST · Top 3, with a RAG learning assistant built and demoed in 24 hours
- 2024
BMW Innovation Challenge · Selected participant, 24-hour challenge at BMW iFactory Dingolfing (DocCheck use case)
- 2020
IEEE Machine Learning Hackathon · 1st place
- 2020
HackCovid-19 · Winner among 130 teams
- 2017
Smart India Hackathon · Winner (Ministry of Defence)
Contact
Open to AI engineering roles in Germany from October 2026. EU Blue Card holder, authorised to work in Germany.