DEV Community

jidonglab profile picture

jidonglab

1 project a week. Building and sharing the entire process — from idea to shipped product in 7 days. Currently: AI news automation.

Why Repetition Penalty Breaks JSON and Code Generation

Why Repetition Penalty Breaks JSON and Code Generation

Comments
7 min read

Want to connect with jidonglab?

Create an account to connect with jidonglab. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
RoPE Scaling: Why Raising rope_theta Breaks Short Context

RoPE Scaling: Why Raising rope_theta Breaks Short Context

Comments
7 min read
Claude Prompt Caching: Why cache_read_input_tokens Stays 0

Claude Prompt Caching: Why cache_read_input_tokens Stays 0

Comments
7 min read
Attention Sinks: Why Evicting Token 0 Wrecks Sliding-Window KV

Attention Sinks: Why Evicting Token 0 Wrecks Sliding-Window KV

Comments
7 min read
Why temperature=0 Still Gives Different Answers: Batch Invariance

Why temperature=0 Still Gives Different Answers: Batch Invariance

Comments
7 min read
Speculative Decoding: Why 80% Acceptance Still Loses at Batch 64

Speculative Decoding: Why 80% Acceptance Still Loses at Batch 64

Comments
8 min read
Why JSON Schema Field Order Breaks Structured Output Accuracy

Why JSON Schema Field Order Breaks Structured Output Accuracy

Comments
7 min read
Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill

Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill

Comments 1
7 min read
DPO Likelihood Displacement: Why Chosen Responses Get Rarer

DPO Likelihood Displacement: Why Chosen Responses Get Rarer

Comments
7 min read
ColBERT Late Interaction: Why MaxSim Beats One Vector Per Chunk

ColBERT Late Interaction: Why MaxSim Beats One Vector Per Chunk

Comments
7 min read
Surface Form Competition: Why Log-Prob Answer Scoring Fails

Surface Form Competition: Why Log-Prob Answer Scoring Fails

Comments
7 min read
Activation Outliers: Why W8A8 INT8 Quantization Needs SmoothQuant

Activation Outliers: Why W8A8 INT8 Quantization Needs SmoothQuant

Comments
8 min read
Cross-Encoder Reranker Score Calibration: Why 0.5 Cutoffs Fail

Cross-Encoder Reranker Score Calibration: Why 0.5 Cutoffs Fail

Comments
8 min read
MCP Tool Sprawl: Why 40 Tools Wreck Tool Selection Accuracy

MCP Tool Sprawl: Why 40 Tools Wreck Tool Selection Accuracy

Comments
7 min read
Why I measure interview silence with WebAudio, not the Speech API

Why I measure interview silence with WebAudio, not the Speech API

1
Comments
5 min read
Best-of-N Sampling: Why N=64 Scores Higher and Answers Worse

Best-of-N Sampling: Why N=64 Scores Higher and Answers Worse

1
Comments
7 min read
Building an AI mock interviewer that tells you why you'd fail: what the turn engine taught me

Building an AI mock interviewer that tells you why you'd fail: what the turn engine taught me

Comments
5 min read
Reciprocal Rank Fusion: Why k=60 Buries Your Best Hit

Reciprocal Rank Fusion: Why k=60 Buries Your Best Hit

Comments
7 min read
Hard Negative Mining Breaks Embedding Fine-Tunes Without Denoising

Hard Negative Mining Breaks Embedding Fine-Tunes Without Denoising

Comments
7 min read
LLM Eval Noise: Why Your 2-Point Accuracy Win Isn't Real

LLM Eval Noise: Why Your 2-Point Accuracy Win Isn't Real

Comments
7 min read
Grouped Query Attention: Why 8 KV Heads Decide Your Batch Size

Grouped Query Attention: Why 8 KV Heads Decide Your Batch Size

Comments
7 min read
Binary Quantized Embeddings: 32x Smaller, If You Center First

Binary Quantized Embeddings: 32x Smaller, If You Center First

Comments
8 min read
MoE Capacity Factor: Why Mixture-of-Experts Drops Your Tokens

MoE Capacity Factor: Why Mixture-of-Experts Drops Your Tokens

Comments
8 min read
Hubness in Vector Search: Why One Chunk Tops Every RAG Query

Hubness in Vector Search: Why One Chunk Tops Every RAG Query

Comments
7 min read
Min-p Sampling: Why top_p Truncates the Wrong Tail

Min-p Sampling: Why top_p Truncates the Wrong Tail

Comments
6 min read
Chunked Prefill: Why One Long Prompt Stalls Every Decode

Chunked Prefill: Why One Long Prompt Stalls Every Decode

Comments
8 min read
Sequence Packing Leaks Across Documents Unless You Mask It

Sequence Packing Leaks Across Documents Unless You Mask It

Comments
6 min read
KV Cache Quantization: Why Keys and Values Need Different Axes

KV Cache Quantization: Why Keys and Values Need Different Axes

Comments
6 min read
Multi-LoRA Serving: Why 100 Adapters Fit on One GPU

Multi-LoRA Serving: Why 100 Adapters Fit on One GPU

Comments
7 min read
Token Healing: Why a Trailing Space Wrecks LLM Completions

Token Healing: Why a Trailing Space Wrecks LLM Completions

Comments
7 min read
Digit Tokenization: Why Commas Fix LLM Arithmetic

Digit Tokenization: Why Commas Fix LLM Arithmetic

Comments
8 min read
Matryoshka Embeddings: Truncate Vector Dimensions, Keep Recall

Matryoshka Embeddings: Truncate Vector Dimensions, Keep Recall

Comments
7 min read
Why Filtered Vector Search Quietly Destroys HNSW Recall

Why Filtered Vector Search Quietly Destroys HNSW Recall

Comments
7 min read
LLM-as-Judge Position Bias: Measure It Before You Ship

LLM-as-Judge Position Bias: Measure It Before You Ship

Comments
8 min read
YaRN vs NTK-Aware RoPE Scaling: Why Long Context Breaks

YaRN vs NTK-Aware RoPE Scaling: Why Long Context Breaks

Comments
7 min read
Prompt Caching: How One Dynamic Token Kills the 90% Discount

Prompt Caching: How One Dynamic Token Kills the 90% Discount

Comments
7 min read
Attention Sinks: Why Streaming LLMs Break When You Evict Token 0

Attention Sinks: Why Streaming LLMs Break When You Evict Token 0

Comments
6 min read
Why Temperature 0 Doesn't Make Your LLM Deterministic

Why Temperature 0 Doesn't Make Your LLM Deterministic

Comments
6 min read
Speculative Decoding: Why a Great Draft Model Still Caps Speedup

Speculative Decoding: Why a Great Draft Model Still Caps Speedup

Comments
6 min read
Constrained Decoding: Force Valid JSON Without Wrecking Accuracy

Constrained Decoding: Force Valid JSON Without Wrecking Accuracy

Comments
6 min read
GRPO Explained: Why DeepSeek Dropped the Critic in RLHF

GRPO Explained: Why DeepSeek Dropped the Critic in RLHF

Comments
6 min read
Online Softmax: How FlashAttention Skips the N N Matrix

Online Softmax: How FlashAttention Skips the N N Matrix

Comments
7 min read
Late Interaction Retrieval: Why ColBERT Beats Single-Vector RAG

Late Interaction Retrieval: Why ColBERT Beats Single-Vector RAG

Comments
7 min read
Why repetition_penalty Quietly Corrupts Your Code Generation

Why repetition_penalty Quietly Corrupts Your Code Generation

Comments
6 min read
DPO Likelihood Displacement: When Preferred Answers Get Rarer

DPO Likelihood Displacement: When Preferred Answers Get Rarer

Comments
6 min read
Why Token Logprobs Beat Asking Your LLM How Confident It Is

Why Token Logprobs Beat Asking Your LLM How Confident It Is

Comments
7 min read
Multi-Token Prediction: DeepSeek's Built-In Draft Model

Multi-Token Prediction: DeepSeek's Built-In Draft Model

Comments
6 min read
AWQ: How Activation-Aware Quantization Saves 4-bit LLMs

AWQ: How Activation-Aware Quantization Saves 4-bit LLMs

Comments
6 min read
PagedAttention: Why Static Batching Wastes Your KV Cache

PagedAttention: Why Static Batching Wastes Your KV Cache

Comments
6 min read
11 Claude Agents Audited My Kit. 2 'Major Bugs' Were Fake

11 Claude Agents Audited My Kit. 2 'Major Bugs' Were Fake

Comments 1
6 min read
Semantic Entropy: Detect LLM Hallucinations Without Ground Truth

Semantic Entropy: Detect LLM Hallucinations Without Ground Truth

Comments
6 min read
Contextual Retrieval: Fix the RAG Chunk That Lost Its Context

Contextual Retrieval: Fix the RAG Chunk That Lost Its Context

Comments
6 min read
Prefilling Claude's Response: Steer Output Without JSON Mode

Prefilling Claude's Response: Steer Output Without JSON Mode

Comments
6 min read
I Cut My Claude Code Prompt Overhead 84% Per Turn (v1.4)

I Cut My Claude Code Prompt Overhead 84% Per Turn (v1.4)

Comments
7 min read
Multi-head Latent Attention: The KV Cache Trick Beyond GQA

Multi-head Latent Attention: The KV Cache Trick Beyond GQA

Comments
7 min read
Why Your Mixture-of-Experts Model Silently Drops Tokens

Why Your Mixture-of-Experts Model Silently Drops Tokens

Comments
7 min read
Embedding Anisotropy: Why Cosine Similarity Never Hits Zero

Embedding Anisotropy: Why Cosine Similarity Never Hits Zero

Comments
6 min read
Return Claude's Thinking Blocks or Your Agent Breaks

Return Claude's Thinking Blocks or Your Agent Breaks

Comments
6 min read
Binary Quantized Embeddings: 32x Smaller Vectors, Recall Intact

Binary Quantized Embeddings: 32x Smaller Vectors, Recall Intact

Comments
7 min read
Prefill/Decode Disaggregation: Stop Serving LLMs on One GPU

Prefill/Decode Disaggregation: Stop Serving LLMs on One GPU

Comments
6 min read
loading...