This project is ongoing. Descriptions, methods and results may be updated after further validation.

Educational and informational content. It is not legal advice, guidance for a specific case or an institutional position. Examples and analyses use only legitimate sources and public data, aggregated or properly anonymised.

Research question

Are drafts produced by retrieval-augmented language models legally reliable?

Motivation

Courts and law firms already use language models to draft decisions. Before relying on them, one needs to measure, with a replicable protocol, where they are right, where they fail and how stable their answers are.

Data

  • Set of test cases and answer keys built with human legal review (being defined)

Methodology

  • Benchmark of language models with and without retrieval-augmented generation
  • Assessment across seven dimensions - legal correctness, reasoning, normative currency, citations, hallucinations, consistency and stability

Limitations

  • Project being structured; there is no public repository or results yet.

Implications

  • Objective criteria for policies on the use of generative AI in the judiciary.

Planned applied products (possible outreach, subject to registration)

  • Public protocol for evaluating AI-generated legal drafts