Related Experiment Video
Updated: Jul 12, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
SteuerLLM: local specialized large language model for German tax law analysis
Sebastian Wind1,2,3, Jeta Sopa4, Laurin Schmid4,5
1Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Martensstr. 3, 91058, Erlangen, Germany. sebastian.wind@fau.de.
Large language models (LLMs) struggle with rule-based domains like tax law. A new domain-adapted LLM, SteuerLLM, trained on authentic German tax law exams, shows superior performance, outperforming larger general models.
Area of Science:
- Artificial Intelligence
- Legal Technology
- Computational Law
Background:
- Large language models (LLMs) excel at general reasoning but falter in rule-intensive domains like tax law.
- Tax law requires precise statutory citation, structured argumentation, and numerical accuracy, posing significant challenges for current AI.
- Existing benchmarks do not adequately capture the complexities of legal examination settings.
Purpose of the Study:
- To introduce SteuerEx, the first open benchmark for German tax law examinations.
- To develop and evaluate SteuerLLM, a domain-adapted LLM specifically for German tax law.
- To demonstrate the effectiveness of domain-specific adaptation over sheer model size for legal AI tasks.
Main Methods:
- Curated SteuerEx, an open benchmark from authentic German university tax-law examinations.
- Developed a statement-level decomposition pipeline for processing examination questions.
- Trained SteuerLLM (28B parameters) on a synthetic dataset using retrieval-augmented generation from authentic materials.
- Implemented a statement-level, partial-credit evaluation framework mirroring examination practices.
Main Results:
- SteuerEx comprises 115 expert-validated questions across six tax law domains and multiple academic levels.
- SteuerLLM consistently outperformed general-purpose LLMs of comparable and larger sizes on tax law reasoning tasks.
- Domain-specific data and architectural adaptation proved more critical than parameter count for performance.
Conclusions:
- Domain-specific LLMs, like SteuerLLM, can achieve high performance in complex legal domains.
- The proposed benchmark and model advance reproducible research in legal artificial intelligence.
- Openly releasing benchmark data, training datasets, and model weights facilitates future advancements in AI for law.
Related Concept Videos
Kohlraush’s Law and its Applications
Lenz's Law
If a bar magnet is moved toward a coil such that the magnetic flux through the coil...
Hess's Law
Actuarial Approach
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...
Modern Molecular Taxonomy
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...