Portrait of Hrishikesh Jadhav

Hrishikesh Jadhav

AI Engineer

AI engineer in Germany with 4+ years shipping production AI and backend systems. I build LLM extraction pipelines, RAG and MCP systems, and the evaluation gates that decide what ships.

I'm Hrishi, and at Toku I built the LLM extraction pipeline, RAG assistant and MCP server that run in daily production inside a global payroll platform. Before that I was a data scientist at GfK / NIQ in Nuremberg, building LLM classification, semantic retrieval and prediction models and running them in production with Python, AWS, Docker and GitLab CI/CD.

Work

  1. Toku help centre

    Customer-facing help centre and AI assistant for payroll, benefits, policy and country guides.

    Help centre retrieval. Two separate retrieval collections, so immigration questions never resolve against payroll content.Customer questionGeneral help corpuspayroll, benefits,policy andcountry guidesVisa support corpusimmigration andwork authorisationLLM answerfallback across model providers
    Two separate retrieval collections, so immigration questions never resolve against payroll content.
    How it works
    Problem
    Give customers one place for payroll, benefits, policy and country-guide answers, with an assistant that answers from that content.
    Architecture
    A Next.js help centre built on Toku's Notion knowledge base. Its assistant retrieves from two separate collections, general help and visa support, with fallback across model providers.
    Hard decision
    Retrieval is split into two collections so that immigration and work-authorisation questions never resolve against general payroll content.
    Evidence
    Helped clients get answers from Toku's Notion knowledge base.
  2. Toku status page

    Public status page for Toku's payroll platform.

    Status page monitoring. Probing the edge and the origin separately tells a CDN failure apart from an application failure, and code fixes go straight to an agent.When an incident opensProbes from the Cloudflare edgeon a short fixed intervalEdge (CDN)Origin (application)Incidents open and resolveautomatically, withdependency rollupsstatus.toku.comJSON API and RSSSlack alertteam notified promptlyCoding agentpicks up failuresneeding a code fix
    Probing the edge and the origin separately tells a CDN failure apart from an application failure, and code fixes go straight to an agent.
    How it works
    Problem
    Catch failures on the payroll platform, hand the ones that need a code fix to an agent automatically, and notify the team on Slack promptly.
    Architecture
    Monitored components are probed from the Cloudflare edge on a short fixed interval. Incidents are detected and resolved automatically, dependencies roll up into one view, and status is published as a JSON API and an RSS feed. When an incident opens, the team is notified on Slack, and failures that need a code fix are picked up by a coding agent.
    Hard decision
    The edge and the origin are probed separately, so a failure is localised to the CDN layer or the application before anyone looks at it.

Experience

  1. 10/202509/2026

    AI Engineer, Toku

    employed via WorkCo Germany GmbH · Frankfurt / Remote

    Global employer-of-record and payroll infrastructure platform operating across 46 countries.

    • Shipped AI tooling into daily production across product, engineering and sales teams: an MCP server exposing internal systems to LLM clients, a RAG assistant (FastAPI, LangChain, pgvector), and n8n automations for onboarding, payslip processing and support triage. Owned deployment, documentation, onboarding and hands-on user support.
    • Cut payroll data intake from hours of manual work per cycle to minutes across production records by building a self-hosted LLM extraction pipeline (Python, FastAPI, PostgreSQL, DigitalOcean) that masks personal data before inference.Self-hosted so that no employee data leaves controlled infrastructure.
    • Kept field-level extraction accuracy above the agreed threshold across every extracted field, checked against a large golden evaluation set, by building the evaluation harness and a CI regression gate that blocks any deploy below it.So a change that lowers extraction accuracy cannot reach payroll data.
    • Moved security review to pull-request time across thousands of pull requests by integrating automated LLM-assisted security review into GitHub CI for Toku's financial platform.
    • Built status.toku.com, Toku's public status platform: production components probed from the Cloudflare edge on a short fixed interval, automated incident detection and resolution, dependency rollups, edge-vs-origin failure isolation, and a JSON API and RSS feed. New incidents notify the team on Slack, and failures that need a code fix are picked up by a coding agent.Probing the edge and the origin separately localises a failure to the CDN or the application before anyone looks at it.
    • Built Toku's customer-facing help centre and AI assistant (toku.com/help, Next.js) on Toku's Notion knowledge base, covering payroll, benefits, policy and country-guide articles across dozens of countries, with general-help and visa-support content in separate retrieval corpora and multi-model fallback across providers.Separate corpora so immigration and work-authorisation questions never resolve against general payroll content.
    • Brought LLM inference costs down to a fraction of the estimated cost of equivalent third-party API usage by deploying open-weight models on self-hosted GPU infrastructure.
    Payroll extraction pipeline. Personal data is masked before inference, and a golden-set gate decides what deploys.In productionBefore every deployPayroll data intakereplaces hours of manual work per cycleMask personal databefore inferenceField extractionself-hosted LLM pipelineGolden evaluation setCI regression gateblocks any deploy belowthe accuracy threshold
    Personal data is masked before inference, and a golden-set gate decides what deploys.
  2. 08/202209/2025

    Data Scientist, GfK - An NIQ Company

    Nuremberg

    Working Student, Data Science, 08/202207/2024

    • Reduced manual effort in product-taxonomy classification by around 30% by building an LLM classification service on Amazon Bedrock using prompt engineering and RegEx-based post-processing.
    • Reduced catalogue lookup time by around 40% versus keyword search by building vector-embedding semantic retrieval over internal product catalogues.
    • Improved simultaneous-viewer prediction accuracy by around 20% over the incumbent model by training and deploying a CatBoost model using 30+ engineered features from sociodemographics, temporal patterns and programme metadata.
    • Kept production scoring pipelines running 24/7 without manual intervention by orchestrating S3 ingestion, feature generation, model scoring and automated integration tests in GitLab CI/CD, containerised with Docker on Linux.
    • Built a TV-audience data-fusion pipeline using K-Nearest Neighbors and a genetic algorithm to match and clone panel households against large-scale return-path data.
  3. 04/202008/2020

    Software Developer (Internship), Sapio Analytics Pvt. Ltd.

    Mumbai

    • Built COVID-19 decision-support models in Python (SEIRD, scikit-learn, SciPy) and forecasting dashboards (Plotly, AWS) used to inform Government of India lockdown and testing-strategy decisions, with a 4.8% RMSE reduction against the prior baseline.

Research

  1. 2025

    Ontology Evolution in Invasion Biology Using Large Language Models: A Hybrid Approach

    Hrishikesh Jadhav, Tina Heger, Birgitta König-Ries, Alsayed Algergawy. LLM-TEXT2KG 2025, CEUR Workshop Proceedings Vol. 4020, pp. 195–206.

    A hybrid pipeline that combines GPT-4 prompting and zero-shot extraction with classical ontology engineering to build and evolve INBIO, a core ontology for invasion biology, with domain experts validating new classes.

    PDF · Proceedings · INBIO ontology

  2. 2021

    A Deep Learning Mobile Application based Sign Language Recognition for Aphasic Person

    Hrishikesh Jadhav, Pushkar Dounde, Akash Pawar, Abhishek Muthange. Journal of Emerging Technologies and Innovative Research (JETIR).

    An Android app that recognises sign-language gestures using Histogram of Oriented Gradients features with CNN and multiclass SVM classifiers.

    Paper · PDF

Awards

  • 2024

    NIQ/GfK HACKFEST · Top 3, with a RAG learning assistant built and demoed in 24 hours

  • 2024

    BMW Innovation Challenge · Selected participant, 24-hour challenge at BMW iFactory Dingolfing (DocCheck use case)

  • 2020

    IEEE Machine Learning Hackathon · 1st place

  • 2020

    HackCovid-19 · Winner among 130 teams

  • 2017

    Smart India Hackathon · Winner (Ministry of Defence)

Contact

Open to AI engineering roles in Germany from October 2026. EU Blue Card holder, authorised to work in Germany.

knowhrishi.de@gmail.com

GitHub · LinkedIn · CV