Aperture's picture

9 1

Aperture

ApertureQA

·

AI & ML interests

None yet

Recent Activity

commented on a paper 1 day ago

SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning

commented on a paper 1 day ago

SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning

replied to umarbutler's post 1 day ago

@abdurrahmanbutler and I just dropped Legal RAG Bench, the first benchmark for legal RAG systems to simultaneously evaluate hallucinations, retrieval failures, and reasoning errors. Our key takeaways are: 1. Embedding models, not generative models, are the primary driver of RAG accuracy. Switching from a general-purpose embedder like OpenAI's Text Embedding 3 Large to a legal domain embedder like Isaacus' Kanon 2 Embedder can raise accuracy by ~19 points. 2. Hallucinations are often triggered by retrieval failures. Fix your retrieval stack, and, in most cases, you end up fixing hallucinations. 3. Once you have a solid legal retrieval engine like Kanon 2 Embedder, it doesn’t matter as much what generative model you use; GPT-5.2 and Gemini 3.1 Pro perform relatively similarly, with Gemini 3.1 Pro achieving slightly better accuracy at the cost of more hallucinations. 4. Google's latest LLM, Gemini 3.1 Pro, is actually a bit worse than its predecessor at legal RAG, achieving 79.3% accuracy instead of 80.3%. These findings confirm what we already knew at Isaacus: that information retrieval sets the ceiling on the accuracy of legal RAG systems. It doesn’t matter how smart you are; you aren’t going to magically know what the penalty is for speeding in California without access to an up-to-date copy of the California Vehicle Code. Even still, to our knowledge, we’re the first to actually show this empirically. Unfortunately, as we highlight in our write-up, high-quality open legal benchmarks like Legal RAG Bench and our earlier MLEB are few and far between. In the interests of transparency, we have not only detailed exactly how we built Legal RAG Bench, but we’ve also released all of our data openly on Hugging Face. You can read our write up [here](https://isaacus.com/blog/legal-rag-bench), noting that we’ll soon be publishing it as a paper. Kudos to my brother @abdurrahmanbutler for serving as the lead author on this monumental release.

View all activity

Organizations

commented 2 papers 1 day ago

SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning

Paper • 2602.13515 • Published 8 days ago • 34 •

SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning

Paper • 2602.13515 • Published 8 days ago • 34 •

replied to umarbutler's post 1 day ago

replied to umarbutler's post 1 day ago

commet

commented on Train AI models with Unsloth and Hugging Face Jobs for FREE 1 day ago

commented on Train AI models with Unsloth and Hugging Face Jobs for FREE 1 day ago

commented on Train AI models with Unsloth and Hugging Face Jobs for FREE 1 day ago

postingonediting

updated a dataset 1 day ago

ApertureQA/renamin

Updated 1 day ago

updated a collection 1 day ago

testing

0 items • Updated 1 day ago

published a Space 1 day ago

Testspacefinal

New activity in ApertureQA/renamin 1 day ago

Create fin

#3 opened 1 day ago by

test_pr

#2 opened 1 day ago by

dataset_discussion

#1 opened 1 day ago by

New activity in ApertureQA/testdataset-hdup 1 day ago

testing

#2 opened 1 day ago by

Create test

#1 opened 1 day ago by

updated a dataset 1 day ago

ApertureQA/testdataset-hdup

Viewer • Updated 1 day ago • 335 • 5

published a dataset 1 day ago

ApertureQA/testdataset-h

Updated 1 day ago

New activity in ApertureQA/renamemodel 1 day ago

testprfinal

#3 opened 1 day ago by

testdiscussion

#2 opened 1 day ago by

testpr

#1 opened 1 day ago by