DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Your RAG copilot can't count — stop letting it try

Highlights how retrieval limits ruin math

Your RAG copilot can't count — stop letting it try

6
Comments 5
6 min read
Hardening an AI coding agent: the failures, and the code that fixed them

Real-world agent failures and loop fixes

Hardening an AI coding agent: the failures, and the code that fixed them

4
Comments 7
27 min read
Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill

Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill

Comments
7 min read
The Requests library for AI one Unified Python SDK for every LLM provider

The Requests library for AI one Unified Python SDK for every LLM provider

Comments
7 min read
GPT-5.6 Luna à 1,40 $/M : on a migré une pipeline de classification, voici la facture

GPT-5.6 Luna à 1,40 $/M : on a migré une pipeline de classification, voici la facture

Comments
5 min read
Skills as Sub-Agents: Orchestrating Complex work with Claude Skills

Skills as Sub-Agents: Orchestrating Complex work with Claude Skills

Comments
7 min read
A Few More Months with Claude — Some Scattered Thoughts (Bite-size Article)

A Few More Months with Claude — Some Scattered Thoughts (Bite-size Article)

Comments
4 min read
DeepSeek V4 Flash vs V4 Pro: A Developer Decision Guide

DeepSeek V4 Flash vs V4 Pro: A Developer Decision Guide

Comments
2 min read
Run DeepSeek V4 Flash 0731 on Your Own Hardware: What It Takes

Run DeepSeek V4 Flash 0731 on Your Own Hardware: What It Takes

Comments
2 min read
Your Agent's Plan Isn't a Plan. It's a Post-Hoc Rationalization

Your Agent's Plan Isn't a Plan. It's a Post-Hoc Rationalization

Comments
5 min read
Real Plugins Need Motors: Skills Should Teach Tools, Not Pretend to Be Them

Real Plugins Need Motors: Skills Should Teach Tools, Not Pretend to Be Them

1
Comments 2
7 min read
Driving vs. delegating: I timed two ways of pairing with AI

Driving vs. delegating: I timed two ways of pairing with AI

Comments
8 min read
The Missing Layer Between LLMs and Reality

The Missing Layer Between LLMs and Reality

Comments
4 min read
Running LLMs Locally on Consumer Hardware — Part 1: The Stack and First Benchmarks

Running LLMs Locally on Consumer Hardware — Part 1: The Stack and First Benchmarks

Comments
3 min read
Anthropic's Opus 5 Release Is About Production Engineering, Not Just Performance

Anthropic's Opus 5 Release Is About Production Engineering, Not Just Performance

Comments 1
2 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.