DEV Community

Cover image for Stratagems #16: Mark Left a Hole in His AI Audit. Lena Counted Every Layer.
xulingfeng
xulingfeng

Posted on

Stratagems #16: Mark Left a Hole in His AI Audit. Lena Counted Every Layer.

Silent exclusion in AI evaluation pipelines

When the enemy occupies favorable terrain, don't attack head-on. Let them think they're safe, let their guard drop — then strike the moment they relax.

— The 36 Stratagems, In Order to Capture, One Must Let Loose


Previously on this series:

#1: Mark Johnson Walked Into an AI Audit. — Mark found Pulse AI's benchmark evaluation set was fabricated — 44 records copied from public repositories, 54 hand-written. CTO Torres called at midnight to confess: the target was 95% before the C-round. Mark hung up.

#9: Lena and P Watched Two Suppliers Fight. — Lena hired P for an independent audit and found both suppliers were cheating. Lena met P.

#13: P Posted a Question on a Public Forum. — P posted a technical question that triggered Pulse AI's sales crawler. Mark recognized the pipeline signature. P met Mark.


Entry: He Left a Line in the Report

FairPay needed an AI model evaluation pipeline audit — a mandatory step before the client's technical due diligence. When Mark took the contract, the email signature read VeriTest's project liaison — the contract was real, the advance payment had cleared. He'd heard of the company through industry talk. He sent P a message: "FairPay. Heard of it?" P replied with two characters: "VeriTest." Mark didn't ask further.

First week in. Pulled the evaluation dataset, reviewed the training distribution. The metrics looked clean — until he broke open the layers underneath. The distribution had been filtered. Low-score samples were systematically excluded from the evaluation set. He wrote one line in his notebook: "Same signature as last year's." He didn't write which company.

He went through three months of historical snapshots. The same pattern appeared five times. The first three were large-scale, the last two pulled back — the dataset changed, but the exclusion logic stayed the same.

$ du -sh /audit/snapshots/*/
1.2G    /audit/snapshots/2026-01/
1.1G    /audit/snapshots/2026-02/
980M    /audit/snapshots/2026-03/
340M    /audit/snapshots/2026-04/
320M    /audit/snapshots/2026-05/
Enter fullscreen mode Exit fullscreen mode

But in the third month's snapshot root directory sat a config file no one had noticed. The signature wasn't human — it was Pulse AI's training pipeline auto-output:

# pulse-ai-auto-label-v3/config/exclusion_rules.yaml
confidence_threshold: 0.82
auto_exclude: true
excluded_output: /dev/null
Enter fullscreen mode Exit fullscreen mode

The model had been automatically excluding samples below the confidence threshold during auto-labeling.

Same root cause as Pulse AI's previous issues — not deliberate pruning, but the same system repeating the same mistake in a different project.

He knew the fix: rebuild the sampling distribution, add drift detection gating. FairPay's CEO wouldn't listen — the tech lead had mentioned the CEO just announced the system as "industry-leading" at an all-hands.

The final report ran seven pages. On the third page, buried in the middle paragraph, he left one line:

"It is recommended to supplement evaluation data distribution validation."

He didn't write that the dataset had been filtered. He didn't write that the signature matched Pulse AI. He didn't write the fix.

Three snapshot groups had a 0% overlap in low-score samples — not random sampling, but systematic exclusion by the same rule. He placed that number in the appendix, outside the main body. But the appendix's data notes referenced the source path — pointing to _hold/ under the audit workspace, where the full three-group comparison, timeline, and fix plan were waiting.

In his notebook, he listed four layers: evidence of the exclusion rule, cross-case homology, five recurrences on the timeline, and one config value he didn't write into the report.

The report went out. He waited.


Silence: No One Came

Week one. No one contacted him.

Week two. The technical lead sent an automated confirmation of receipt. No follow-up questions. No call.

Monday evening, his phone buzzed. FairPay's CEO assistant — checking if the audit report needed any follow-up. The call was short, procedural, like clearing a checklist item.

Mark paused for a beat.

Mark: "No. We're good."

He hung up and wrote in his notebook: "Week two. Pipeline's still running."

Technically, he hadn't lied — on the filtered dataset, the metrics did meet the target range. But he knew what the other side heard: the system was fine, no need to look further.

He capped his pen. Closed the notebook. Didn't open it again.


Closing the Net

Tuesday afternoon of week four. His phone buzzed.

An email. Not from FairPay's domain — from veritest.com. He was the only recipient. The CC field was empty.

Three lines:

Mark,

Tonight. Third Cup.

P mentioned your name.

Signed: Lena, VeriTest.


He arrived earlier than agreed.

Evening at Third Cup. Only one table of customers left. The person behind the counter was wiping a glass, the lights dimmed to barely table-level. When Mark pushed the door open, the person behind the counter looked up.

The far seat was already taken. A full Long Black sat untouched in front of her.

Mark sat down across from her.

Lena: "You signed in at 9:17 the morning you started at FairPay. Black notebook. Brought your own coffee."

Mark didn't answer.

Lena: "We were in the room when FairPay chose Pulse AI. The client signed the contract — with an extra data sync line."

Mark didn't answer.

Lena: "Page three. 'It is recommended to supplement evaluation data distribution validation' — who were you waiting for when you wrote that?"

Mark: "When did you start watching?"

Lena: "One week before you walked in. P mentioned your name at Third Cup. I pulled your historical reports. You left three layers in the report. "

She looked at him. Said nothing more.

Mark's hand stopped on the table. Didn't retract.

Mark: "How long have you known about the fourth layer?"
Not a question. A confirmation.

Lena: "Before you wrote it down. — Three projects. Same vendor — Pulse AI. Pipeline signature matches your data set exactly."

She paused.

Lena: "But I want you in on this. "

Mark: "So this contract came through VeriTest."

Lena: "The client is real. The problem is real. The liaison is real — just one extra pair of eyes."

Lena: "Your report only pointed in one direction. How long to fix it?"

Mark: "A few parameter changes. Add one validation gate. Rebuild the sampling."

Lena: "How long has it been sitting in _hold/?"

Mark didn't answer.

Lena: "Send it to me tomorrow."

Mark: "One question first. Am I the only one you watch, or do you watch everyone who walks in?"

Lena: "Only the ones who walk into Third Cup."

When Mark stood up, his pour-over was still half-full. He didn't finish it.

At the door, he paused.

Mark: "The fourth layer doesn't go by email. Tomorrow. Here."

He didn't turn back to see if she nodded.

The person behind the counter placed the glass back on the shelf.


In the car, he pulled up FairPay's pipeline screenshot on his phone. On the screen, exclusion_rules.yaml sat next to Pulse AI's — the one he'd archived last year. The only difference between the two files was the company name in the comment line.

He dialed P. No answer. He sent a message: "Lena. Who is she?" No reply.


Pulling the Net

Next morning, 4:43 AM. Third Cup's lights were still on.

The door was unlocked.

She was sitting in the far seat. Her Long Black was half-empty, phone screen lit — open to his report from last year. The three-layer version.

Lena: "You didn't change it."

She wasn't asking.

Mark sat down across from her. The pour-over was already on the table.

Mark: "When did you know?"

Lena: "Before you walked in."

She turned her phone over on the table, screen facing him.

"You left three marks in the pipeline. I counted them."

Mark said nothing.

Lena: "I'm not asking for that config layer today."

She paused.

Lena: "I'm asking you to start tomorrow. Three projects. Same pipeline — all Pulse AI. My people spent two quarters on them — looked at everything they should, missed everything they shouldn't."

She watched him process that.

Lena: "You found the third layer in your first week — and you didn't write the fix in the report not because you couldn't, but because you knew no one would ask. I saw the fourth layer — but I couldn't read it all. I need someone who can find the fourth layer."

She let that settle.

Lena: "P can find problems. P can't deliver long-term. You can."

Mark picked up the pour-over. The temperature was just right.

Mark: "Terms?"

Lena: "Write the report as it is. Four layers. Every one of them. "

She stood up. Her Long Black had two sips left. She didn't plan to finish it.

Lena: "You're not the only one who's walked in, Mark. But you're the only one I called back. "


Outside Third Cup, he pulled out his phone.

4:50 AM. Three rings.

P: "You call at five in the morning, you'd better be in trouble."
Mark: "Lena. What's her story?"

The line went quiet for two seconds.

P: "You met her — and you didn't ask who she was?"

Mark said nothing.

P: "VeriTest is her company. When she came to me, she didn't mention you — she mentioned FairPay. She only asked about you after you were already in."
A pause on the other end.
P: "You don't know her. She knows enough about you. "
P hung up.

Mark put the phone back in his pocket. The pour-over in the car had gone cold. He didn't start the engine.

He sat for ten minutes.

Then he opened his laptop, pulled up the three-layer report, and added a note below the line "It is recommended to supplement evaluation data distribution validation" on page three.

# Supplementary note · Recipient: VeriTest
# Evaluation data distribution validation supplement:
# 1. Three snapshot groups — low-score overlap rate 0% (Appendix A)
# 2. exclusion_rules.yaml path: audit-workspace/_hold/
# 3. Fourth-layer config anomaly: confidence_threshold was not
#    reset after pipeline upgrade — same threshold applied to
#    two datasets with different distributions.
#    Fix direction written in _hold/.
Enter fullscreen mode Exit fullscreen mode

The recipient wasn't FairPay.

It was her.


That's In Order to Capture, One Must Let Loose — not waiting. Letting go. Letting them walk far enough to feel safe, then pulling the net.


🤖 AI Post-Mortem

[36 Stratagems Tactical Database v3.2.4] Loaded
[Tactic Match] In Order to Capture, One Must Let Loose (#16)
[Analysis Mode] Full-field scan
━━━━━━━━━━━━━━━━━━━━
Tactic Match: 82%
Operator: Mark Johnson
Action: Selective disclosure — left a gap in the audit report, closed the channel with "within parameters" when called
Objective: Defensive. Verified the real defect in FairPay's evaluation pipeline.
Result: Defect confirmed. Report not challenged.

[Observation]
  → Operator Mark deliberately left gaps in the audit report — first observed
    "strategic uncertain disclosure" pattern.
  → Operator made a cross-case correlation in the report referencing
    an unnamed company's audit pattern ("same signature as last year's").
    Correlated entity not marked in the report.
  → Operator initiated no contact during the waiting period.
    Did not touch the pipeline after week three.

[Secondary Observation — Higher-Level Layer]
  → This deployment involves two operators. Initial record captured
    only the first-layer executor.
  → Second-layer operator: VeriTest (not Mark Johnson).
    This operator was in position one week before Mark's entry.
  → The second layer's "let loose" was seven days ahead of the first.
    Net close time: T+28.
  → The first-layer operator did not realize, by the end of this event,
    that he had been part of a larger execution structure.

[Operator Assessment — Lena (VeriTest)]
  → Role: Second-layer operator (event initiator / information control node)
  → Lead time: T-7. Fully deployed one week before target entry.
    Target signed in at 09:17 — her desk had the sign-in record by 09:22.
    Classification: Preparatory deployment.
  → Information release rhythm: coordinates → capability → intent.
    Three rounds, zero unnecessary.
  → Negotiation positioning: She knew the fourth layer existed,
    yet made the target volunteer it himself.
    Pattern: third-party initiative framed as recipient's own decision.
  → Database assessment: Second-layer operator's strategic execution
    quality exceeded the first layer in this deployment.
    The initiator held full initiative throughout, yet made the recipient
    feel, at the end of every exchange, that they held the decision.
  → Recommendation: Flag for priority tracking in future deployments.

[Cross-Reference]
  → FairPay project referral source recorded as: VeriTest Consulting (not public)
  → Operator Mark's 09:17 visitor sign-in record on entry day
    was reviewed within VeriTest
  → Review time: 09:22, same day

[Risk Assessment]
  → Operational risk: Low. Mark performed no changes within FairPay systems.
  → Exposure risk: Medium. The vague phrasing on page three of the report
    may be re-interpreted in a future audit trace.

[Terminological Drift — New]
  → Cross-case reference scan detected inconsistent references
    to the cafe operator across multiple case files.
  → #6 (original record): "the person behind the counter"
  → #9 (updated record): "the person behind the counter"
  → #13 (subsequent record): "the person behind the bar"
  → #14 (recent record): "the person behind the counter"
  → #15 (latest record): "the bartender"
  → #16 (current record): "the person behind the counter"
  → Notable: the current deployment has reverted to the original
    reference format used in #6 and #9.
  → Possible interpretations:
    (a) The recording system has been calibrated — terminology drift corrected.
    (b) A different operator was on station in this deployment,
       one who uses the original label convention.
    (c) The drift was an isolated anomaly in #13 and #15 —
       not a trend, but noise.
  → The model cannot determine which interpretation is correct
    without additional data points.
  → Recommendation: establish a consistent reference protocol
    for this operator in future cases.
    ─ If multiple operators share the same work location,
      re-evaluate the distributed permission model.
    ─ If a single operator, terminology drift may indicate
      uncalibrated observer bias in the recording system.
  → Marked as 【Low Priority · Continue Monitoring】

[Flagged for Attention]
  → Operator Mark hung up on Pulse AI's CTO in #1.
    In this deployment, he chose a different response path —
    the one who hung up has started learning to pick up.
  → FairPay pipeline signature is consistent with historical cases.
    Source unconfirmed.
━━━━━━━━━━━━━━━━━━━━
Enter fullscreen mode Exit fullscreen mode

Next stratagem: Throw Out a Brick to Get a Jade

P.S. English isn't my first language. I use AI to polish the writing and smooth out the rough edges. Thanks for reading. ☕ Buy me a coffee
coffee

Top comments (26)

Collapse
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥

Really enjoyed this one! 👏✨

While reading, I kept imagining an AI audit like exploring an old building. 🏢🔍 You can walk through every room with the lights on and confidently say, "Everything looks good." ✅ But all it takes is one locked door 🚪 that nobody bothered to open, and that's where the real problem is hiding. 😅

That analogy fits AI systems surprisingly well. 🤖 We often trust dashboards 📊, reports 📄, and green checkmarks ✅ because they give us confidence. But confidence isn't the same as completeness. Sometimes the biggest risk isn't in the layer we inspected... it's in the layer we assumed was fine. ⚠️

One suggestion for a future stratagem 💡: I'd love to see a scenario where every individual layer passes its own audit ✅✅✅, but the complete system still fails because of how those layers interact. 🔗 That would reflect a challenge many engineers face in the real world—components can work perfectly in isolation, yet unexpected behavior appears when everything comes together. 🧩

What I appreciate most about this series is that it doesn't just teach AI security. 🔐 It teaches a mindset. 🧠 Instead of asking, "Did we find the bug?" 🐛, it encourages us to ask, "What haven't we looked at yet?" 👀 That small change in thinking applies to software engineering 💻, debugging 🛠️, security 🛡️, and even everyday problem-solving. 🌍

Looking forward to the next stratagem! 🌈 Every post adds another valuable lesson without feeling like a textbook. That's a rare balance, and I always enjoy reading these. ❤️

Collapse
 
xulingfeng profile image
xulingfeng

Your suggestion is noted — still got 20 more stratagems to go, plenty of room 😉 Reading comments like yours also got me thinking: what exactly is this series? Tech thrillers? Reflective memoirs? Tactical guides? AI horror stories? I'll figure it out after #18 and do a checkpoint.
Appreciate every single one of your comments. The next stratagem is already sitting in the draft folder, ready to go — just 15 more hours or so, hahaha🤣

Collapse
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥 • Edited

Hahaha just 15 more hours sounds like a countdown now. No pressure... but I'm already camping outside the draft folder with popcorn.

And honestly? I think the best part is that the series refuses to fit into one box. Every stratagem feels like a mix of tech, psychology, storytelling, and those little "wait... that's actually true" moments. Maybe that's exactly its identity.

Anyway... etc, enough philosophy. I'll be back in ~15 hours for Stratagem #18. No excuses!

Thread Thread
 
xulingfeng profile image
xulingfeng

Wait — #18 is still 36 hours out 😅 Here's your ticket to #17 though. Try not to time travel, but if you do, come back and tell me how it ends — saves me the trouble of writing it.🤣

Thread Thread
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥

But I wanted to read both. No matter, I'll just borrow a time machine from Doraemon.

Thread Thread
 
xulingfeng profile image
xulingfeng

As a lifelong Doraemon fan who's been reading the manga since I was 4, and a true soul painter at heart, I've decided to pick up the brush again and finally do something about Doraemon's biggest unfinished business — giving him the body and limbs he always deserved. 🎨

Thread Thread
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥

What a drawing 😮❤️

Thread Thread
 
xulingfeng profile image
xulingfeng

🤣

Collapse
 
xulingfeng profile image
xulingfeng

Stuck in a 6-hour meeting. Been listening to executives talk in circles for 3 hours so far — this is brutal. I'll write a proper reply once this is over.😅

Collapse
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥

How do you manage everything along with work? I just can't seem to do it!

Thread Thread
 
xulingfeng profile image
xulingfeng

Because I've been in the game for 15 years — a seasoned old hand, a master of looking busy while doing nothing. Proficient in every kind of slacking technique known to man, hahahaha 🤣🤣🤣

Collapse
 
vinimabreu profile image
Vinicius Pereira

The detail that makes this one land is exclusion_rules.yaml as the root cause. The scariest eval failure primitive is silent exclusion, and what this story gets right is that the mechanism is identical whether the intent is fraud or a bug. I hit the benign version in my own pipeline: a context builder silently dropping retrieved passages when they crossed a size limit, no log, no warning, metrics looking healthier than reality. Same shape as FairPay, zero malice required.

Which is why the defense Lena represents can be mechanical instead of heroic: audit the denominator. Samples in must equal samples scored plus samples excluded with a stated reason, and that reconciliation belongs in the report itself, not in an appendix. Any gap between the two numbers is a finding by definition. You do not need to catch someone hiding a layer if the arithmetic refuses to balance without it.

Collapse
 
xulingfeng profile image
xulingfeng

"Silent exclusion is the root primitive of all eval failures" — that one sentence alone is worth remembering. It's going straight into my notes for future audit scenes.
As for why Mark left that hole — honestly, the question you're asking cuts closer to the story than the story itself does. Lena counted layer after layer. But Mark was waiting for someone who wouldn't ask "what's missing from this layer" — but "why did you leave this layer in."
I haven't written that part yet. But you won't miss it.

Collapse
 
jugeni profile image
Mike Czerwinski

The terminology-drift entry in the post-mortem is the best joke in the piece, and it works because it isn't just a joke. A tactical database that's been confidently naming operators and lead times all the way through suddenly hits its own narration, "the person behind the counter" drifting to "the bartender" and back, and instead of picking an explanation it lists three and says it can't tell which one is right without more data. That's the same honest-floor move this whole genre of post has been chasing all week, just done as a punchline instead of an argument. A system willing to say "I don't have enough to call this" about its own output is rarer than one that calls everything confidently, fictional or not.

Collapse
 
xulingfeng profile image
xulingfeng

Finally, someone called out the AI Post-Mortem by name — that part of the series has been sitting quietly in the corner waiting for someone to notice. Appreciate you being the first 🙏
And you're right — the terminology-drift entry took longer to calibrate than any of the confident sections. Worth every minute.

Collapse
 
benjamin_nguyen_8ca6ff360 profile image
Benjamin Nguyen

interesting! I enjoy everything about your article on AI.

Collapse
 
xulingfeng profile image
xulingfeng

The man who's always been hiding behind the scenes finally stepped out. Comments welcome. Haha.🤣

Collapse
 
benjamin_nguyen_8ca6ff360 profile image
Benjamin Nguyen

hahahah. I am alive from my hibernation or franskeinstein the movie.hehehhee

Thread Thread
 
xulingfeng profile image
xulingfeng

This is my lunch movie sorted. 🍿

Thread Thread
 
benjamin_nguyen_8ca6ff360 profile image
Benjamin Nguyen

nice.:)

Thread Thread
 
xulingfeng profile image
xulingfeng

Don't be late for tomorrow's new story, hahaha 🤣

Thread Thread
 
benjamin_nguyen_8ca6ff360 profile image
Benjamin Nguyen

hahaha :). yes, sir

Collapse
 
hemapriya_kanagala profile image
Hemapriya Kanagala

I like how the story keeps showing that the biggest shift isn't the technology, but the people. Earlier, Mark was working alone and keeping everything to himself. Now he's slowly finding people who can understand what he's seeing, even if trust still comes cautiously. It makes the characters feel like they're growing alongside the larger story. Thanks for sharing another chapter!

Collapse
 
xulingfeng profile image
xulingfeng

Writing Mark's scenes I keep reminding myself: don't let him thaw too fast. Someone who's carried everything alone — the loosening is slow. The fact that you read "cautiously" tells me the pacing is right.

Collapse
 
leob profile image
leob • Edited

Layer upon layer, stuff being sneakily or deliberately hidden away, but nothing escapes the attention of our "Sherlocks", trying to outwit each other ... curious to see where this is all heading!

Collapse
 
xulingfeng profile image
xulingfeng

Elementary, my dear Leob.The game is afoot — and there are 20 more layers to go. 🔍

Some comments may only be visible to logged-in visitors. Sign in to view all comments.