Reference-system proposal · not a partner claim

Join Trainium inference and F2 FPGA memory as one EC2 reference architecture.

EC2 Trn2 inference paired with an F2 HBM FPGA memory service. Engrammatic proposes a bounded test: keep dense model execution on AWS Trainium2 on EC2 Trn2, move durable session state into a compact memory path, and measure how many matched long sessions the same allocation can serve.

Claim boundary: the 8× density and 9.4× aggregate-decode result below was measured by Engrammatic on one NVIDIA L40S—not on Amazon Web Services hardware. The company-specific session density remains intentionally unstated until the reference test. No Amazon Web Services endorsement is implied.
Measured by Engrammatic · Aug 23 202616 vs 2 sessions · 8× density · 9.4× aggregate decode

One 48 GB NVIDIA L40S, a 32B-class model, matched responses, and the same 16-second p95 latency gate.

Reference-system targetMeasure first · price second

No session-density or savings number is asserted for this stack before the matched benchmark.

Amazon Web Services × Engrammatic · mechanism · node · evidence · request1:59 · narration · optional captions
01 · The memory problem

Replay and retained history spend accelerator memory on the past.

Long-running agents either resend a growing transcript or retain an expanding KV cache. Engrammatic changes the memory path, not the model's dense math.

Without Engrammatic

History grows every turn

  • More tokens are replayed or retained.
  • KV cache competes with model weights and active batches.
  • Long sessions push the scheduler toward more accelerator allocations.
With Engrammatic

Working context stays fixed

  • Exact mutable facts are written once.
  • Relationships are encoded into compact 10,240-bit HDC patterns.
  • Only the current working set returns to the model.
02 · Proposed Amazon Web Services node

Four clear jobs. One measurable system.

Trainium2 serves the model. An F2 instance evaluates compact HDC memory, while EC2 networking carries only fixed-size working context and state updates.

Model computeAWS Trainium2 on EC2 Trn2

Dense transformer inference and token generation stay on the existing accelerator path.

Compact recallEC2 F2 · VU47P HBM FPGA

XOR, popcount, and fixed-size retrieval are placed here only after the benchmark confirms the best boundary.

Exact stateEC2 host + EFA

Authoritative counters, identifiers, permissions, tool results, and updates remain exact.

OperationsAWS Neuron / EKS

The current scheduler, runtime, metrics, and customer control surface remain in place.

03 · Business case

Raise useful sessions per accelerator—not an invented GPU-hour price.

A new memory-aware inference pattern that can consume two existing EC2 accelerator families instead of requiring custom silicon first.

Economic unit
sessions / accelerator

The benefit exists only if more simultaneous long sessions pass the same response-quality and service-level gates. We do not assume a node price or claim a savings percentage before that measurement.

Decision rule

Ship only if the whole system wins.

Compare the growing-history baseline and Engrammatic on the same model, prompts, hardware allocation, and traffic. Count transfer, host, memory-card, software, and energy overhead.

response qualityp95 latencytokens / secondenergy / sessionsession density
04 · Evidence and sources

Public facts below. Proposed integration above.

The product facts come from Amazon Web Services's public material. The architecture is an Engrammatic hypothesis that still requires engineering validation.

Company technology

Amazon EC2 F2 instances ↗

EC2 F2 provides AMD-Xilinx VU47P HBM FPGA accelerators; EC2 Trn2 provides Trainium2 acceleration.

Engrammatic evidence

One measured serving comparison

16 constant-memory sessions versus 2 long-history sessions on one 48 GB L40S; same 16-second p95 gate; 8× density and 9.4× aggregate decode. Source package and raw benchmark are available for diligence.

One bounded collaboration

A small Trn2 and F2 allocation, EFA guidance, and a joint same-model service-level benchmark.

Engrammatic will implement the memory adapter, run the matched benchmark, document every assumption, and report a negative result if the quality, latency, power, or density gate fails.

Discuss the reference-system test