Kurdish Speech Logo
Kurdish Speech
← Back to articles
Build agent memory with NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors
Large Language Models & Generative AI

Build agent memory with NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors

In my previous post: Building persistent memory for multi-agent AI systems with Amazon S3 Vectors , we explored why memory engineering is the foundational discipline for production multi-agent systems. We showed how Amazon S3 Vectors , a capability of Amazon Simple Storage Service (Amazon S3), meets the architectural requirements for agent memory: semantic retrieval, rich metadata, strong consistency, and elastic scale. In this post, we move from architecture to implementation. We show you how to use Amazon S3 Vectors as the persistent memory layer within the NVIDIA NeMo Agent Toolkit (NAT) , deployed on Amazon Elastic Kubernetes Service (Amazon EKS) for full operational control. By the end of this post, you will understand how NAT’s memory subsystem works and how to implement Amazon S3 Vectors as a custom memory provider. You will also learn how to deploy the stack on Amazon EKS, using a multi-agent investment research use case as the running example. What is NVIDIA NeMo Agent Toolkit? NVIDIA NeMo Agent Toolkit (NAT) is an open source framework for building, profiling, and optimizing AI agents. It’s framework-agnostic, working with Strands Agents, LangChain, LlamaIndex, CrewAI, and custom implementations. NAT provides four capabilities relevant to production agent systems: Agent orchestration – Define agents as composable workflows with configurable large language models (LLMs), tools, and prompts. You run them locally with nat run or as persistent services with nat serve . Profiling – Track token usage, latency, throughput, and run times across agents and individual tools

Source: AWS Artificial Intelligence

Source: AWS Artificial Intelligence