Steven Gonsalvez

Software Engineer

harveyai/harvey-labs: A benchmark built to evaluate and improve agent capabilities for supporting legal work.

Why CEREBRO kept it

Agent benchmark for legal work directly on-topic

The text below is an automated extraction of the article at https://github.com/harveyai/harvey-labs, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (github.com).

Legal Agent Benchmark (LAB): An open-source benchmark for evaluating agents on real legal work. Harvey LAB is an open-source project aimed at benchmarking LLM agents' abilities to perform legal work in realistic environments. LAB consists of two parts: a dataset of tasks containing agent instructions, documents, and rubrics as well as an execution harness for running and evaluating agents against those tasks. LAB is an ongoing project and we expect to consistently add to and refine the task set and execution harness. Read the announcement post: Introducing Harvey's Legal Agent Benchmark Start

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/agents · ai/llm-mechanics

Also from github.com