Technical Post
RAG in practice: implementing Amazon Bedrock Knowledge Bases with OpenSearch Serverless and Terraform
Large Language Models have radically changed the way companies see search, automation, and interaction with corporate knowledge. The…
Large Language Models have radically changed the way companies see search, automation, and interaction with corporate knowledge. The trouble starts when these applications leave the demo environment and go into production. At that point, it quickly becomes clear that the model does not know the company.

The model does not know internal contracts, operational runbooks, organizational policies, private catalogs, specific procedures, or the latest documentation for the environment. Without access to that context, corporate applications end up producing generic, inconsistent, or simply wrong answers. In many cases, the problem is not just answer quality, but operational risk.
This is exactly the scenario in which Retrieval-Augmented Generation (RAG) has become one of the most relevant architectures in the current generation of enterprise AI applications.
The idea is relatively simple: before generating an answer, the system retrieves relevant information from a private knowledge base and uses that context as grounding for the model. In practice, this turns the LLM from a purely statistical system into a contextualized interface over the organization's knowledge.
In this article, we will build a RAG architecture using:
- Amazon Bedrock
- Amazon Bedrock Knowledge Bases
- Amazon OpenSearch Serverless
- Amazon S3
- Terraform
The goal is not just to provision resources, but to discuss the architectural aspects that typically come up when RAG applications start running in real environments.
The architecture
The proposed architecture has four core components.
Corporate documents are stored in S3. Bedrock Knowledge Bases is responsible for document processing, chunking, embedding generation, and retrieval orchestration. OpenSearch Serverless acts as the vector store. Terraform provisions all the infrastructure.
The operational flow is fairly straightforward:
Usuário
|
Aplicação / API
|
Bedrock Knowledge Base
|
OpenSearch Serverless
|
Foundation Model
|
Resposta contextualizada
Although it looks simple conceptually, most of the operational complexity shows up in the details: chunking, synchronization, IAM, access control, document governance, and observability.
Project structure
The Terraform structure used in this article is deliberately simple.
rag-bedrock-terraform/
├── providers.tf
├── variables.tf
├── s3.tf
├── iam.tf
├── opensearch.tf
├── bedrock.tf
└── outputs.tf
The goal here is not to build a complete enterprise-ready module, but to clearly show the components involved in building the RAG pipeline.
Configuring the provider
The AWS provider you use needs to support the Bedrock Knowledge Bases resources.
terraform {
required_version = ">= 1.6.0"
required_providers {
aws = {
source = "hashicorp/aws"
version = ">= 5.49.0"
}
}
}
provider "aws" {
region = var.aws_region
}
The project variables stay minimal:
variable "aws_region" {
default = "us-east-1"
}
variable "project_name" {
default = "rag-demo"
}
Provisioning the document bucket
The S3 bucket will serve as the document source for the Knowledge Base.
data "aws_caller_identity" "current" {}
resource "aws_s3_bucket" "documents" {
bucket = "${var.project_name}-documents-${data.aws_caller_identity.current.account_id}"
}
resource "aws_s3_bucket_versioning" "documents" {
bucket = aws_s3_bucket.documents.id
versioning_configuration {
status = "Enabled"
}
}
Versioning matters because RAG applications inevitably run into problems with document updates, rollback, and synchronization inconsistencies. In corporate environments, the problem is rarely just storing documents; it is usually keeping multiple versions of organizational knowledge consistent.
Creating the vector store
OpenSearch Serverless will be used as the application's vector database.
Before the collection, we need to define the encryption policy.
resource "aws_opensearchserverless_security_policy" "encryption" {
name = "${var.project_name}-encryption"
type = "encryption"
policy = jsonencode({
Rules = [
{
Resource = [
"collection/${var.project_name}-vectors"
],
ResourceType = "collection"
}
],
AWSOwnedKey = true
})
}
Next, we provision the vector collection.
resource "aws_opensearchserverless_collection" "vectors" {
name = "${var.project_name}-vectors"
type = "VECTORSEARCH"
depends_on = [
aws_opensearchserverless_security_policy.encryption
]
}
The VECTORSEARCH type enables vector search and nearest neighbor search, which are fundamental to semantic retrieval.
This is where an important difference between prototypes and real environments already shows up. In the lab, any vector database seems to work well. In production, latency, cost, throughput, and operational governance quickly start to matter.
Bedrock permissions
Bedrock will need access to:
- the documents in S3;
- the embedding models;
- and OpenSearch Serverless.
The base role looks like this:
resource "aws_iam_role" "bedrock_role" {
name = "${var.project_name}-bedrock-role"
assume_role_policy = jsonencode({
Version = "2012-10-17"
Statement = [
{
Effect = "Allow"
Principal = {
Service = "bedrock.amazonaws.com"
}
Action = "sts:AssumeRole"
}
]
})
}
The bucket access policy:
resource "aws_iam_policy" "s3_policy" {
name = "${var.project_name}-s3-policy"
policy = jsonencode({
Version = "2012-10-17"
Statement = [
{
Effect = "Allow"
Action = [
"s3:GetObject",
"s3:ListBucket"
]
Resource = [
aws_s3_bucket.documents.arn,
"${aws_s3_bucket.documents.arn}/*"
]
}
]
})
}
And the model access policy:
resource "aws_iam_policy" "bedrock_policy" {
name = "${var.project_name}-bedrock-policy"
policy = jsonencode({
Version = "2012-10-17"
Statement = [
{
Effect = "Allow"
Action = [
"bedrock:InvokeModel"
]
Resource = "*"
}
]
})
}
The policies are then attached to the role.
resource "aws_iam_role_policy_attachment" "s3_attach" {
role = aws_iam_role.bedrock_role.name
policy_arn = aws_iam_policy.s3_policy.arn
}
resource "aws_iam_role_policy_attachment" "bedrock_attach" {
role = aws_iam_role.bedrock_role.name
policy_arn = aws_iam_policy.bedrock_policy.arn
}
In many projects, IAM ends up being treated as a secondary detail. In enterprise RAG applications, the opposite usually happens. The problem quickly becomes document access control, knowledge segregation, and permission-aware retrieval.
Creating the Knowledge Base
Now we provision the Knowledge Base.
resource "aws_bedrockagent_knowledge_base" "rag_kb" {
name = "${var.project_name}-kb"
role_arn = aws_iam_role.bedrock_role.arn
knowledge_base_configuration {
type = "VECTOR"
vector_knowledge_base_configuration {
embedding_model_arn = "arn:aws:bedrock:${var.aws_region}::foundation-model/amazon.titan-embed-text-v2:0"
}
}
storage_configuration {
type = "OPENSEARCH_SERVERLESS"
opensearch_serverless_configuration {
collection_arn = aws_opensearchserverless_collection.vectors.arn
vector_index_name = "rag-index"
field_mapping {
vector_field = "vector"
text_field = "text"
metadata_field = "metadata"
}
}
}
}
From here, Bedrock takes over managing:
- embedding generation;
- chunking;
- synchronization;
- retrieval;
- integration with the vector store.
This significantly reduces the initial complexity of the implementation, although it does not eliminate the architectural challenges of RAG.
Connecting the document source
The document source will be the S3 bucket provisioned earlier.
resource "aws_bedrockagent_data_source" "documents" {
knowledge_base_id = aws_bedrockagent_knowledge_base.rag_kb.id
name = "${var.project_name}-documents"
data_source_configuration {
type = "S3"
s3_configuration {
bucket_arn = aws_s3_bucket.documents.arn
}
}
}
After terraform apply, documents can be uploaded to the bucket as usual.
aws s3 cp ./sample-docs \
s3://SEU_BUCKET/ \
--recursive
After that, just start the Knowledge Base sync.
During this process:
- the documents are processed;
- chunks are created;
- embeddings are generated;
- vectors are stored in OpenSearch.
The problem that almost always shows up: chunking
Few aspects affect RAG quality as much as chunking.
In most prototypes, chunking is treated as simple text splitting. In production, it becomes one of the main drivers of semantic retrieval quality.
Chunks that are too large increase cost, degrade precision, and waste context. Chunks that are too small break semantic meaning and reduce coherence.
The problem gets worse with corporate documents, because many of them contain:
- tables;
- hierarchical structures;
- cross-references;
- interdependent sections;
- organizational acronyms;
- specific terminology.
In practice, strategies based purely on size often produce poor results. Semantic chunking usually works better.
RAG does not fix bad documentation
This may be the biggest reality check in corporate projects.
Many organizations implicitly assume that AI will solve:
- inconsistent documentation;
- fragmented knowledge;
- lack of governance;
- duplicated content;
- outdated documentation.
In practice, RAG amplifies precisely the quality of the information that already exists.
If the documents are bad, so will the system be.
Costs and operations
RAG applications introduce new operational components:
- embeddings;
- vector storage;
- retrieval;
- synchronization;
- inference;
- reindexing.
At small scale, this may seem irrelevant. In corporate environments, cost quickly starts to matter, especially for OpenSearch Serverless and token consumption.
On top of that, RAG applications need to be observable.
In production, questions inevitably come up:
- which chunks were retrieved;
- how relevant they were;
- which document the answer came from;
- how much context was sent;
- what the cost per query was;
- what the hallucination rate was.
Without observability, troubleshooting becomes extremely difficult.
Conclusion
RAG has established itself as one of the most important architectures in the current generation of enterprise AI applications. However, implementing RAG in production goes far beyond connecting a model to a vector database.
The main challenges usually show up in:
- document governance;
- chunking;
- IAM;
- synchronization;
- observability;
- cost;
- and the quality of organizational knowledge.
Amazon Bedrock Knowledge Bases significantly reduces the initial operational complexity, while Amazon OpenSearch Serverless provides a managed, scalable vector database.
With Terraform, this entire architecture can be reproduced in a consistent, auditable way, aligned with modern infrastructure-as-code practices.
In the end, though, the biggest problem with RAG is rarely the model.
It is the maturity of the organization's own knowledge.
See you in the next post!
Comments
Every comment is moderated before it appears here. Nothing is published automatically.
Loading…