One 48 GB NVIDIA L40S, a 32B-class model, matched responses, and the same 16-second p95 latency gate.
EC2 Trn2 inference paired with an F2 HBM FPGA memory service. Engrammatic proposes a bounded test: keep dense model execution on AWS Trainium2 on EC2 Trn2, move durable session state into a compact memory path, and measure how many matched long sessions the same allocation can serve.
One 48 GB NVIDIA L40S, a 32B-class model, matched responses, and the same 16-second p95 latency gate.
No session-density or savings number is asserted for this stack before the matched benchmark.
Long-running agents either resend a growing transcript or retain an expanding KV cache. Engrammatic changes the memory path, not the model's dense math.
Trainium2 serves the model. An F2 instance evaluates compact HDC memory, while EC2 networking carries only fixed-size working context and state updates.
Dense transformer inference and token generation stay on the existing accelerator path.
XOR, popcount, and fixed-size retrieval are placed here only after the benchmark confirms the best boundary.
Authoritative counters, identifiers, permissions, tool results, and updates remain exact.
The current scheduler, runtime, metrics, and customer control surface remain in place.
A new memory-aware inference pattern that can consume two existing EC2 accelerator families instead of requiring custom silicon first.
The benefit exists only if more simultaneous long sessions pass the same response-quality and service-level gates. We do not assume a node price or claim a savings percentage before that measurement.
Compare the growing-history baseline and Engrammatic on the same model, prompts, hardware allocation, and traffic. Count transfer, host, memory-card, software, and energy overhead.
The product facts come from Amazon Web Services's public material. The architecture is an Engrammatic hypothesis that still requires engineering validation.
EC2 F2 provides AMD-Xilinx VU47P HBM FPGA accelerators; EC2 Trn2 provides Trainium2 acceleration.
16 constant-memory sessions versus 2 long-history sessions on one 48 GB L40S; same 16-second p95 gate; 8× density and 9.4× aggregate decode. Source package and raw benchmark are available for diligence.
Engrammatic will implement the memory adapter, run the matched benchmark, document every assumption, and report a negative result if the quality, latency, power, or density gate fails.
Discuss the reference-system test