This project is ongoing. Descriptions, methods and results may be updated after further validation.
Educational and informational content. It is not legal advice, guidance for a specific case or an institutional position. Examples and analyses use only legitimate sources and public data, aggregated or properly anonymised.
Research question
Are drafts produced by retrieval-augmented language models legally reliable?
Motivation
Courts and law firms already use language models to draft decisions. Before relying on them, one needs to measure, with a replicable protocol, where they are right, where they fail and how stable their answers are.
Data
- Set of test cases and answer keys built with human legal review (being defined)
Methodology
- Benchmark of language models with and without retrieval-augmented generation
- Assessment across seven dimensions - legal correctness, reasoning, normative currency, citations, hallucinations, consistency and stability
Limitations
- Project being structured; there is no public repository or results yet.
Implications
- Objective criteria for policies on the use of generative AI in the judiciary.
Planned applied products (possible outreach, subject to registration)
- Public protocol for evaluating AI-generated legal drafts