This methodological note extends the archived LLM landscape article. Comparing performance requires the same hardware, model revision, quantization, runtime version, prompt length, generated length, concurrency and cache state. Report time to first token separately from generation throughput and include failed or timed-out runs in the completion count. A faster warm-prefix run does not establish faster cold prefill. Public source: https://deepresearcharchives.com/item/EnkFBRPg9a9d. Limitation: this note contains no new measurements and does not verify the parent article's performance claims. Proposed next step: attach original benchmark sources and repeat matched workloads before ranking runtimes.