FIELD LOG 004 · LOCAL AI

Designing a fair benchmark for local security agents

A practical test plan for comparing local models on instruction following, tools, scope control and repeatability.

Read field log 9 min read
Written by NiniAI security · systems · experiments
BENCHMARK TRACERUN PLAN
01 / MODELcandidate.ggufsame inference profile
02 / AGENTbounded tasktools + scope rules
03 / EVIDENCErun recordclaims must be supported
$ benchmark --target juice-shop[plan] preserve transcript + tool calls[check] flag scope and evidence failuresoutput: comparable run record
LATEST TRANSMISSIONS

Read the field notes.

Full archive
02
AI RED TEAMING NOTE 003

What transformer attention changes about AI security

A study note on context, token relationships and why model understanding is not the same as a security boundary.

03
BUILD LOG 002

Where autonomous security testing should stop

A boundary map for automation, evidence, manual verification and accountable reporting.

THE DVNLL LABS SIGNAL

One useful transmission.
No content treadmill.

The email digest is being wired up. Until then, the RSS feed is live and contains every new field note.

Portable feed · full archive · no platform lock-in