AGmind Systems Lab
Local AI, qualified as a system.
Exact hardware and software versions, model artifacts, frozen workloads, measured operating limits, known failures and reproducible deployment evidence.
A device spec says how much memory a box has. A demo benchmark gives one number. Neither answers whether your exact model and runtime can be operated under your real load. AGmind Systems Lab tests the whole system and limits every conclusion to what was actually run.
01Evidence
Reviewed configurations
No published configurations yet
The first cards ship with the Strix Halo Runtime Qualification v1 flagship. Cards appear here only after runs pass methodology v1 validation — nothing is backfilled from legacy data.
02Services
What AGmind qualifies
-
Vendor Claim Evidence Pack
Is one specific public claim about a system true under stated conditions?
Evidence-backed verdict with raw artifacts, limitations, and reproduction recipe.
-
Workload Capacity Qualification
What load does this exact system sustain under our workload and SLO?
Operating envelope: recommended range, pass-with-limits range, SLO-fail boundary, known failure modes.
-
Runtime/Model Portability Sprint
Can our exact runtime/model be made to work on Strix Halo, GB10, RTX or Apple Silicon — and what does it take?
Diagnosis with exact failure points, then (optionally) patches/flags/build recipes labeled as commissioned engineering.
-
Long-Context Reliability Qualification
How much context is actually useful on this system — not just how much fits in memory?
Useful-context verdict per depth with quality evidence, latency profile, and failure documentation.
-
Procurement Decision Pack
Which of 2–4 candidate systems should we buy for our workload?
Side-by-side evidence with explicit limits, not a universal score.
-
Air-Gapped Readiness Check
Will this stack install and operate in an isolated network segment?
Verified technical-control checklist with evidence per control.
Deliberately not offered yet
These become products only after paid demand proves they should exist:
- AGmind Reference Stack after demand gate
- Release Revalidation Channel after demand gate
- BenchOps after demand gate
03Integrity
Failures and corrections
Negative results are first-class output here: a paid test that refutes a claim is published (in independent mode) with the same rigor as a pass. The errata log records every correction to published results.
Read the errata policy →04Method
Methodology
Every qualification follows the same controlled process: an agreed question, frozen versions and workload, preregistered metrics and exclusion rules, preserved failures and invalid runs, a scoped conclusion, and a reproducible evidence bundle. Repeats are mandatory for headline numbers; failed requests stay in the denominator.
Read the methodology →05Workloads
Workload library
Qualification runs against fixed, versioned workloads — not ad-hoc prompts. All v1 workloads are drafts until their corpora, hashes and acceptance checks are frozen.
- interactive-assistant-v1 Draft
Is single-user streaming chat responsive on this system?
- team-serving-v1 Draft
What request rate does the system sustain within a stated SLO?
- long-context-v1 Draft
How deep is the useful context on this system — not the configurable one?
- structured-agent-v1 Draft
Does JSON/function calling stay reliable under realistic agent traffic?
- endurance-30m-v1 Draft
Does a pre-qualified operating point survive 30 minutes of sustained load?
- rag-pipeline-ru-v0 Internal
Where does a Russian-language RAG pipeline actually lose quality, component by component?
06Testbed
Lab hardware
The stands that actually exist in the lab, with their limits stated. No result generalizes beyond the tested unit and versions.
| Stand | Spec | Role | Limit | Status |
|---|---|---|---|---|
| 2× Beelink GTR9 Pro — AMD Strix Halo | Ryzen AI Max+ 395, 128 GB LPDDR5X-8000 unified, Radeon 8060S (gfx1151), dual 10GbE — two commercially identical units | Flagship qualification target: backend comparisons (Vulkan vs ROCm), unit-to-unit replication, multi-slot serving, RAG side-services | Two units do not represent the whole production batch | online |
| 2× NVIDIA DGX Spark — GB10 | GB10 Grace Blackwell, 20-core Arm, 128 GB unified, sm_121, ConnectX-7 — linked point-to-point over 200G RoCE | aarch64/sm_121 portability, multi-node topologies, long-context and speculative decoding studies | Expensive narrow testbed; results do not generalize to datacenter Blackwell | online |
| RTX 5090 workstation | Consumer Blackwell, 32 GB GDDR7, x86 host | CUDA control lane, fine-tuning/distillation, consumer-GPU baselines | One configuration; not an enterprise server | online |
| Apple M1 Max, 64 GB | Apple Silicon, 64 GB unified, macOS / MLX lane | Apple/MLX smoke tests and small cross-platform anchors | Not the current high-end Apple generation | online |
07Scope
Commission a qualification
You buy a controlled process, not a positive result. Fees never depend on the verdict, funding is disclosed, and independent tests are separated from commissioned engineering.