Search

6 results for "RAG"

Models & Releases

Comparing Continued Pretraining to RAG (accuracy and performance)

Mostly as a fun experiment I wanted to do a quick comparison of performance and accuracy between a CPT trained QWEN 3.5 4B model and a RAG implementation against the base model. The point of this exercise is mostly to measure the perform...

R/ r/LocalLLaMA
Models & Releases

Looking for advice for small office looking for local AI RAG

My company is looking for basically a local hardware back-end for an already set up Open WebUI Windows AD joined setup. We are already on the frontier models thru Open WebUI. ~10 users. Workload is a lot of contract and chat based email...

R/ r/LocalLLaMA
Models & Releases

Expert Support case study: Bolstering a RAG app with LLM-as-a-Judge

HU Hugging Face Blog
Models & Releases

Building Cost-Efficient Enterprise RAG applications with Intel Gaudi 2 and Intel Xeon

HU Hugging Face Blog
Models & Releases

Anthropic spent this week in hot water over cybersecurity

After admitting earlier this year that its AI models had hacked other companies' systems on a handful of occasions, Anthropic released a new report on Wednesday detailing the attacks. It reveals a string of incidents displaying what Anth...

TH The Verge AI
Micron's memory wall chart. Compute up ~3x every two years, HBM bandwidth under 2x
Models & Releases

Micron's memory wall chart. Compute up ~3x every two years, HBM bandwidth under 2x

From Raghu Sreeramaneni's memory tutorial at hot chips 2026. top line is normalised tflops for tpu v3 through r200, bottom line is hbm2e through hbm4, both log scale, so the distance between them is a lot wider than it looks. The three b...

R/ r/LocalLLaMA