DEV Community

Patrick Hughes profile picture

Patrick Hughes

404 bio not found

Joined Joined on  github website
VRAM Fit Is Not Runtime Support

VRAM Fit Is Not Runtime Support

Comments
4 min read

Want to connect with Patrick Hughes?

Create an account to connect with Patrick Hughes. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
Why Local LLM Benchmarks Need Power Data

Why Local LLM Benchmarks Need Power Data

Comments
4 min read
My 5090 benchmark was missing the field I needed most

My 5090 benchmark was missing the field I needed most

Comments
4 min read
Search Old Results Before Publishing an LLM Test

Search Old Results Before Publishing an LLM Test

Comments
5 min read
Build Local LLM Eval Data From Real Failures

Build Local LLM Eval Data From Real Failures

Comments
5 min read
How I Keep LLM Results Valid After a Driver Update

How I Keep LLM Results Valid After a Driver Update

Comments
4 min read
How I Budget VRAM for Shared Local AI Workloads

How I Budget VRAM for Shared Local AI Workloads

Comments
5 min read
Why I Benchmark Local LLM Input and Output Separately

Why I Benchmark Local LLM Input and Output Separately

Comments
5 min read
Why I Test 3 Workloads Before Sizing a Local LLM

Why I Test 3 Workloads Before Sizing a Local LLM

Comments
5 min read
Incident response needs a local model you already trust

Incident response needs a local model you already trust

Comments
6 min read
Local open-model agents just became a product category

Local open-model agents just became a product category

Comments
6 min read
Your local LLM benchmark is probably lying to you

Your local LLM benchmark is probably lying to you

Comments
5 min read
Test Retrieval Before Your Local LLM Writes

Test Retrieval Before Your Local LLM Writes

Comments
4 min read
Why I Did Not Promote My Smaller Local Model

Why I Did Not Promote My Smaller Local Model

Comments
4 min read
Build an AI Research Workbench on Your Own GPU

Build an AI Research Workbench on Your Own GPU

Comments
5 min read
Ollama num_batch: 256 Was My RTX 5090 Sweet Spot

Ollama num_batch: 256 Was My RTX 5090 Sweet Spot

Comments
4 min read
A 7 GB 27B Model Lost to My 17 GB Default

A 7 GB 27B Model Lost to My 17 GB Default

1
Comments
5 min read
My Agents Have to Prove What They Did

My Agents Have to Prove What They Did

Comments 1
4 min read
Your Agent's Audit Trail Cannot Be Retrofitted

Your Agent's Audit Trail Cannot Be Retrofitted

Comments
5 min read
Ollama Raised $65M. What Builders Get

Ollama Raised $65M. What Builders Get

1
Comments
5 min read
Why production AI is moving to open weights

Why production AI is moving to open weights

Comments
4 min read
Devlog 2026-07-13: local drive git corruption stalls the q

Devlog 2026-07-13: local drive git corruption stalls the q

Comments
3 min read
My 8B Model Failed a 400-Word Task

My 8B Model Failed a 400-Word Task

Comments
5 min read
Q4km vs Q5km: Q4_K_M vs Q5_K_M

Q4km vs Q5km: Q4_K_M vs Q5_K_M

Comments
3 min read
I Rebuilt 5,383 Embeddings After a Dimension Change

I Rebuilt 5,383 Embeddings After a Dimension Change

Comments
5 min read
I Raised the QA Bar on Blog Infographics

I Raised the QA Bar on Blog Infographics

Comments
4 min read
Pin Your Context Window or Pay the Reload Tax

Pin Your Context Window or Pay the Reload Tax

Comments
5 min read
Preview Retrieval Before Your Local LLM Runs

Preview Retrieval Before Your Local LLM Runs

Comments
5 min read
Pin Your Local LLM Context Size Before You Build a Router

Pin Your Local LLM Context Size Before You Build a Router

Comments
4 min read
When Claude hits a weekly limit, your agent fleet still needs a third CLI

When Claude hits a weekly limit, your agent fleet still needs a third CLI

Comments
4 min read
I built a self-improving code model on one RTX 5090. Here is what actually worked.

I built a self-improving code model on one RTX 5090. Here is what actually worked.

Comments
5 min read
Will That Local Model Fit? Do the VRAM Math First

Will That Local Model Fit? Do the VRAM Math First

Comments
5 min read
Local LLMs Need a Timeout Before They Need a Bigger Model

Local LLMs Need a Timeout Before They Need a Bigger Model

Comments
5 min read
Your local LLM is not a worse Claude. It is a different tool.

Your local LLM is not a worse Claude. It is a different tool.

Comments
5 min read
How I Gate a Local Coding Model Before I Trust It

How I Gate a Local Coding Model Before I Trust It

Comments
5 min read
A verifier loop beats a faster local model

A verifier loop beats a faster local model

Comments
4 min read
How I Make Local Model Runs Fail Safely On A 5090

How I Make Local Model Runs Fail Safely On A 5090

Comments
5 min read
How to Make a Local QLoRA Starter Fail Safely

How to Make a Local QLoRA Starter Fail Safely

Comments
5 min read
Use Owner Gates and AgentGuard to Keep AI Agents Moving

Use Owner Gates and AgentGuard to Keep AI Agents Moving

Comments
5 min read
How to Run Local LLM Verifier Loops on Owned Hardware

How to Run Local LLM Verifier Loops on Owned Hardware

Comments
5 min read
AI Agent Memory: What Actually Works in 2026

AI Agent Memory: What Actually Works in 2026

Comments
4 min read
What Anthropic's MITRE ATT&CK report means for solo AI builders

What Anthropic's MITRE ATT&CK report means for solo AI builders

Comments
4 min read
AI Agent Memory in 2026: How It Works and When to Use It

AI Agent Memory in 2026: How It Works and When to Use It

Comments 1
2 min read
Your AI Agent Says "Done." Make It Prove It.

Your AI Agent Says "Done." Make It Prove It.

2
Comments 2
4 min read
Give Your AI Agents an Append-Only Event Log

Give Your AI Agents an Append-Only Event Log

1
Comments
4 min read
A self-healing system can't heal an empty queue

A self-healing system can't heal an empty queue

Comments
4 min read
Missing AI agent cost data is not zero

Missing AI agent cost data is not zero

Comments 1
4 min read
Anthropic Writes 80% of Its Code with Claude

Anthropic Writes 80% of Its Code with Claude

Comments
2 min read
What Salesforce's 20,000 AI Agent Deployments Teach a Solo Builder

What Salesforce's 20,000 AI Agent Deployments Teach a Solo Builder

Comments 2
4 min read
57-71% of AI agents leak data between users. Here's what to do.

57-71% of AI agents leak data between users. Here's what to do.

Comments
3 min read
VRAM Calculator: Estimate Local LLM Requirements

VRAM Calculator: Estimate Local LLM Requirements

Comments
1 min read
Anthropic's IPO and the 40% Cost-Savings Gap: Why Your Spend Cap Matters More Now

Anthropic's IPO and the 40% Cost-Savings Gap: Why Your Spend Cap Matters More Now

Comments
4 min read
When JPMorgan's AI bill goes up, who controls it?

When JPMorgan's AI bill goes up, who controls it?

Comments
4 min read
57-71% of AI Agents Leak Data Between Users. Here's the Fix.

57-71% of AI Agents Leak Data Between Users. Here's the Fix.

Comments
4 min read
AI Coding Assistant Pricing in 2026: Copilot vs Cursor vs Claude Code

AI Coding Assistant Pricing in 2026: Copilot vs Cursor vs Claude Code

Comments
3 min read
How to Close the AI Agent Cost Gap at the Call Site

How to Close the AI Agent Cost Gap at the Call Site

Comments
4 min read
Agentic coding moved my bottleneck to code review

Agentic coding moved my bottleneck to code review

Comments
3 min read
How to Pick a GGUF Quant Level for Your VRAM Budget

How to Pick a GGUF Quant Level for Your VRAM Budget

Comments
4 min read
Stop Telling People You Have 11 AI Agents

Stop Telling People You Have 11 AI Agents

Comments
4 min read
GGUF Quantization and VRAM: How to Pick Q4, Q5, or Q8 for Your GPU (2026)

GGUF Quantization and VRAM: How to Pick Q4, Q5, or Q8 for Your GPU (2026)

Comments
4 min read
loading...