Talento

LLMOps200 pool

LLM evaluation framework for claims workflows

You will build or extend an LLM evaluation system that benchmarks a multi-step agentic claims-triage workflow against a set of test claims (with ground-truth routing/outcomes). Use Langfuse, OpenLit, or a custom harness to measure accuracy, latency, cost, and detect regressions when models or prompts change. Provide a dashboard or report showing which agents/tools contribute most to failures.

  • Test dataset of 10–20 realistic insurance claims with expected routing/handling decisions
  • Evaluation harness that runs the workflow on each claim and compares output to ground truth
  • Metrics: accuracy, precision/recall per routing category, token cost, latency, and tool success rate
  • Integration with Langfuse or OpenLit to log traces and surface failure patterns
  • Summary report identifying which agents, tools, or prompts are most likely to cause misrouting

Skills & tools

PythonCrewAI or LangGraphLangfuse or OpenLitTesting & metrics designSQL or data analysis (for failure analysis)

Context

You will build or extend an LLM evaluation system that benchmarks a multi-step agentic claims-triage workflow against a set of test claims (with ground-truth routing/outcomes). Use Langfuse, OpenLit, or a custom harness to measure accuracy, latency, cost, and detect regressions when models or prompts change. Provide a dashboard or report showing which agents/tools contribute most to failures.

Deliverables

  • Test dataset of 10–20 realistic insurance claims with expected routing/handling decisions
  • Evaluation harness that runs the workflow on each claim and compares output to ground truth
  • Metrics: accuracy, precision/recall per routing category, token cost, latency, and tool success rate
  • Integration with Langfuse or OpenLit to log traces and surface failure patterns
  • Summary report identifying which agents, tools, or prompts are most likely to cause misrouting

Success metrics

Prize pool of €200 — split between top submissions (60 / 25 / 15%) · this task is open to contract hire — strong submissions may lead to a longer-term role

Codebase access Tier A — No codebase access

Ready to apply?

Connect your GitHub to build a profile and apply for this task in seconds.

Sign in with GitHub