
An open-source benchmark for evaluating long-term memory in AI agents across multi-session conversations, including recall, temporal reasoning, fact updates, speaker attribution, refusal, credibility, and tool-use tasks.
Prices and medians update for the tier you select.
Ranked by how closely each one matches FP-AMB's job. Prices show each provider's Pro state; entry prices are labelled as such. Unpriced products still belong to the market.
Market = the products most similar to this one by capability; prices are median / quartiles over its priced members, separated by provider type and buyer tier.