This project builds a fully serverless, event-driven lead routing system on AWS for approximately £0/month at small agency scale. A lead arrives via HTTP POST. Within 3 seconds it is stored in DynamoDB, the right agent is notified by email, and the event is durably queued for asynchronous processing. No human intervention. No inbox. No delay.
The system implements two parallel processing paths intentionally: a synchronous path guarantees immediate storage and notification before the HTTP response is returned, while an asynchronous path through EventBridge and SQS guarantees durability — if the processor fails, the message is retained for up to 4 days and retried automatically. No lead is ever lost.
All 19 AWS resources are provisioned by Terraform in under 2 minutes and destroyed completely with terraform destroy. AWS resources are deployed in the eu-central-1 region.
I spent two years at Foxtons, one of London's largest estate agencies. Every day I watched the same thing happen: a lead came in from Rightmove, a website form, or Facebook Ads — and it sat in someone's inbox for hours before an agent picked it up. By the time they called back, the buyer had already spoken to three other agencies.
The fix exists. Salesforce Essentials costs £25 per user per month. For a 6-person agency that's £150/month minimum, plus setup costs, plus training. Most small agencies don't buy it. So leads keep getting lost.
This project builds that fix on AWS for approximately £0/month at small agency scale — fully automated, fully serverless, zero servers to manage.
A lead arrives. Within 3 seconds it is stored, the right agent is notified, and the pipeline records the event. No human intervention. No inbox. No delay.
The system intentionally processes each lead via two parallel paths:
Synchronous path (lead-handler → DynamoDB + SNS): guarantees immediate storage and instant agent notification before the HTTP response is returned. Sub-100ms.
Asynchronous path (EventBridge → SQS → lead-processor): guarantees durability. If lead-processor fails, the message remains in SQS for up to 4 days and retries automatically. No lead is ever lost.
This pattern trades a small amount of duplication for both low latency and high reliability — a deliberate design decision, not an oversight.
Stack: Python 3.11 · boto3 · Terraform · AWS Lambda · API Gateway · DynamoDB · SNS · SQS · EventBridge · CloudWatch · IAM
Region: eu-central-1 · Resources deployed: 19 · Clean destroy: confirmed
Lead Source (HTTP POST)
│
▼
API Gateway ── POST /lead endpoint
│
▼
Lambda: lead-handler ── Python 3.11 · 29ms avg execution
│
├──▶ DynamoDB (leads-table) ── Immediate storage · UUID partition key
│
├──▶ SNS (lead-notifications) ── Instant agent email notification
│
└──▶ EventBridge (lead-bus) ── Async event emission
│
▼
EventBridge Rule (lead-rule)
│
▼
SQS (lead-queue) ── 4-day retention · SSE-SQS encrypted
│
▼
Lambda: lead-processor ── SQS consumer · batch size 1
│
├──▶ DynamoDB (leads-table) ── Durable write confirmation
│
└──▶ SNS (lead-notifications) ── Pipeline confirmation email
CloudWatch ── Logs + Metrics + lambda-errors-alarm
IAM ── Shared execution role · least-privilege per service
| Service | Role |
|---|---|
| Amazon API Gateway | Exposes the public HTTP POST /lead endpoint |
| AWS Lambda (lead-handler) | Ingests lead, writes to DynamoDB, publishes to SNS, fires EventBridge event |
| AWS Lambda (lead-processor) | Consumes from SQS, writes to DynamoDB, publishes pipeline confirmation via SNS |
| Amazon DynamoDB | Stores all lead records — on-demand capacity, UUID partition key |
| Amazon SNS | Delivers real-time email notifications to confirmed agent subscribers |
| Amazon SQS | Durable message buffer between EventBridge and lead-processor — 4-day retention |
| Amazon EventBridge | Custom event bus decoupling lead-handler from downstream consumers |
| Amazon CloudWatch | Logs, metrics dashboard, and lambda-errors-alarm across both functions |
| AWS IAM | Least-privilege execution role granting only required permissions per service |
| Terraform | Provisions and destroys all 19 resources as code |
- A lead is submitted via HTTP POST to the API Gateway endpoint with
name,email, andmessagefields - API Gateway proxies the request to the lead-handler Lambda function
- lead-handler generates a UUID
lead_id, writes the record to DynamoDB, and publishes an SNS notification — all before returning the HTTP 200 response - lead-handler fires a
LeadCreatedevent to the custom EventBridge bus with the lead data in the eventdetail - The EventBridge rule matches the event on
source: lead.apiand routes it to SQS - SQS durably queues the message — if lead-processor is unavailable, the message is retained for up to 4 days and retried automatically
- lead-processor is triggered by the SQS event source mapping, unwraps the EventBridge envelope from the SQS record body, and writes a confirmation record to DynamoDB
- lead-processor publishes a pipeline confirmation notification via SNS
- CloudWatch collects invocation metrics, duration, and error counts from both Lambda functions
- The lambda-errors-alarm triggers if errors exceed 1 per minute across either function
aws-property-lead-router/
│
├── architecture/
│ └── 00-aws-property-lead-router-architecture.png # End-to-end architecture overview
│
├── lambda/
│ ├── handler.py # lead-handler: ingest → DynamoDB + SNS + EventBridge
│ └── processor.py # lead-processor: SQS consumer → DynamoDB + SNS
│
├── terraform/
│ ├── main.tf # All 19 AWS resource definitions
│ └── outputs.tf # API Gateway endpoint URL
│
├── screenshots/
│ ├── 01-project-structure.png
│ ├── 02-terraform-init.png
│ ├── 03-dynamodb-table-active.png
│ ├── 04-sns-topic-created.png
│ ├── 05-sns-subscription-confirmed.png
│ ├── 06-lambda-handler-deployed.png
│ ├── 07-api-gateway-route.png
│ ├── 08-api-gateway-terraform-apply.png
│ ├── 09-api-test-curl-success.png
│ ├── 10-dynamodb-leads-stored.png
│ ├── 11-email-notification-received.png
│ ├── 12-cloudwatch-handler-logs.png
│ ├── 13-sqs-queue-created.png
│ ├── 14-lambda-processor-deployed.png
│ ├── 15-sqs-trigger-enabled.png
│ ├── 16-eventbridge-rule-active.png
│ ├── 17-dynamodb-full-pipeline-6-items.png
│ ├── 18-email-sqs-pipeline-confirmed.png
│ ├── 19-cloudwatch-metrics-dashboard.png
│ ├── 20-cloudwatch-alarm-configured.png
│ └── 21-terraform-destroy-complete.png
│
├── .gitignore
└── README.md
- AWS CLI configured with appropriate credentials
- Terraform >= 1.0 installed
- Docker (optional — for local Lambda testing)
terraform init # Download hashicorp/aws v5.100.0 provider
terraform plan # Preview all 19 resources before apply
terraform apply # Provision full stack (~2 minutes)After apply, confirm the SNS email subscription by clicking the link in the AWS confirmation email sent to your agent address. The stack is not fully operational until the subscription is confirmed.
curl -X POST https://<your-api-id>.execute-api.eu-central-1.amazonaws.com/lead \
-H "Content-Type: application/json" \
-d '{"name":"Test Lead","email":"test@email.com","message":"Interested in a 2-bed flat"}'Expected response: HTTP 200 with a confirmation message. Within 3 seconds, the lead appears in DynamoDB and an email notification arrives at the confirmed subscriber address.
terraform destroy # Remove all 19 resources · zero idle cost after teardownhashicorp/aws v5.100.0 installed and locked. Provider lock file committed for reproducible deployments.
leads-table provisioned via Terraform. Partition key: lead_id (UUID). Capacity mode: on-demand. Status: Active.
lead-notifications topic with confirmed email subscription. Agent notification pipeline active.
Email subscription confirmed via AWS confirmation link — notifications now delivered to the agent inbox.
lead-handler deployed — Python 3.11. Writes to DynamoDB, publishes to SNS, fires to EventBridge in a single invocation.
POST /lead route configured on lead-api with Lambda integration wired.
curl -X POST with real lead data returning 200 OK and confirmation message.
4 lead records in leads-table after initial testing — all with UUID partition keys and ISO 8601 timestamps.
SNS → Gmail confirmed. Real agent notifications delivered within seconds of the API call.
lead-handler execution — Duration: 29.49ms, Billed: 479ms, Memory: 128MB used: 87MB. Zero errors.
lead-queue — Standard type, SSE-SQS encryption, 4-day message retention.
lead-processor deployed — Python 3.11. SQS event source mapping configured. Consumes from lead-queue, writes to DynamoDB, publishes to SNS.
SQS event source mapping active: lead-queue → lead-processor. State: Enabled. Batch size: 1.
lead-rule on lead-bus — Status: Enabled. Routes events from source: lead.api into lead-queue.
6 items after full pipeline test — including entries confirming the EventBridge → SQS → lead-processor → DynamoDB path works end-to-end.
Pipeline confirmation email triggered by the full async path — EventBridge → SQS → lead-processor → SNS → Gmail working.
Lambda metrics — Invocations, Duration, Errors, Throttles across both functions in a single view.
lambda-errors-alarm — Errors > 1 for 1 datapoint within 1 minute. Status: Insufficient data (correct — no errors fired).
terraform destroy — all 19 resources destroyed. Zero orphaned infrastructure. Zero idle cost.
Why Lambda over EC2 or Fargate? Lead routing is event-driven and spiky — a form is submitted, a function runs for 29ms, then it is idle. Lambda is the correct compute model. Fargate would be correct for a persistent web application. EC2 for sustained predictable load. The access pattern here is pure event-driven: Lambda wins on cost, simplicity, and zero operational overhead.
Why DynamoDB over RDS? The access pattern is simple: write a lead record, read it by lead_id. No joins. No relational queries. DynamoDB on-demand at this volume costs cents per month. RDS has a fixed monthly baseline cost — the wrong tool for a key-value access pattern.
Why EventBridge instead of direct Lambda-to-SQS? EventBridge decouples the producer (lead-handler) from downstream consumers. The handler emits an event without knowing what processes it. Rules route events to SQS, which acts as a durable buffer. This allows adding new consumers later — analytics, CRM integrations, audit pipelines — without modifying handler.py. Direct Lambda-to-SQS would tightly couple ingestion to a single consumer.
Why SQS between EventBridge and lead-processor? SQS provides durability that direct Lambda invocation cannot. If lead-processor errors, the message stays in the queue and retries automatically for up to 4 days. Zero leads lost even if the processor is down for hours. This is the reliability layer of the system.
Why a custom EventBridge bus? Using lead-bus instead of the default AWS event bus isolates application events from AWS service events in the same account. Cleaner observability, easier to scale, and future event patterns are less likely to conflict with AWS-generated events.
Why Terraform over console provisioning?
Every resource is defined in code. terraform apply provisions the full stack in under 2 minutes. terraform destroy removes all 19 resources completely. No manual steps, no configuration drift, no orphaned resources. The destroy step is as important as the apply — if infrastructure cannot be cleanly removed, it is not production-ready.
Why a shared IAM role? Both Lambda functions share a single execution role for simplicity in this project. In production, separate roles per function would enforce strict least-privilege. Documented in Production Improvements.
Why no VPC, IGW, NAT, ALB, or subnets? All services used are AWS-managed and require no VPC placement. Lambda has native outbound internet access. Adding a VPC would cost ~£30/month for a NAT Gateway and introduce latency with zero security benefit for this architecture. The absence of VPC is a deliberate design decision, not an omission.
After deploying the SQS event source mapping and confirming it was enabled in the console, lead-processor was not being invoked when messages arrived in the queue. The SQS messages were visible in the console and the queue depth was growing, but Lambda was making no attempt to consume them. No errors appeared in the Lambda console — the function was simply silent.
The root cause was a missing IAM permission. The execution role for lead-processor did not include sqs:ReceiveMessage or sqs:DeleteMessage. Lambda requires both permissions to poll a queue and acknowledge processed messages. Without them, the event source mapping is enabled but functionally inert — Lambda cannot pull from the queue and produces no visible error at the function level. The failure was only visible in CloudWatch Logs as an access denied exception on the polling attempt.
Fix: Added explicit sqs:ReceiveMessage, sqs:DeleteMessage, and sqs:GetQueueAttributes to the IAM role policy in main.tf and re-ran terraform apply. The event source mapping began delivering messages to lead-processor immediately after the policy update.
Lesson: IAM permissions must match the exact API calls a service makes — not just the service name. Lambda polling SQS requires three distinct permissions. Always verify the CloudWatch Logs for the function itself when an event source mapping appears enabled but the function is not triggering — the access denied exception is logged there, not in the Lambda console.
With the EventBridge rule active and confirmed enabled, messages published to lead-bus were not arriving in lead-queue. The SQS queue remained empty after every API call, the lead-processor was never invoked, and no error appeared anywhere in the console. The EventBridge rule showed as enabled with the correct target configured.
The cause was a missing SQS resource-based policy. EventBridge requires explicit sqs:SendMessage permission granted on the target queue via a queue policy — this is separate from the Lambda execution role and is not inherited from any other IAM configuration. Without it, EventBridge attempts delivery and silently discards every message. This was the hardest failure to diagnose in the entire project because no error was surfaced anywhere in the EventBridge or SQS consoles.
Fix: Added an aws_sqs_queue_policy Terraform resource explicitly granting EventBridge sqs:SendMessage on lead-queue, scoped to the EventBridge service principal and the specific rule ARN. Messages routed correctly to the queue on the next apply.
Lesson: EventBridge → SQS always requires an explicit queue resource policy. It is not inherited from any IAM role. Silent failure with no visible error is the only symptom. Test the full event routing path end-to-end immediately after provisioning by sending a real event and verifying the SQS queue depth increases — do not assume the configuration is correct based on console status alone.
After the EventBridge → SQS → Lambda path was working end-to-end, lead records were appearing in DynamoDB from lead-processor but all fields were empty — name, email, and message were None for every record written by the async path. The synchronous path via lead-handler was writing correctly; the issue was isolated to lead-processor.
The cause was an incorrect assumption about the payload structure. When EventBridge routes an event to SQS, the original JSON body is wrapped inside an EventBridge envelope. The SQS record body is not the raw lead data — it is the full EventBridge event object with the lead data nested under a detail key. The initial processor.py was parsing record['body'] and accessing name, email, and message directly, which returned None for all fields because they don't exist at the top level of the envelope.
Fix: Updated processor.py to unwrap the envelope before accessing lead fields:
event_body = json.loads(record['body'])
body = event_body["detail"] # unwrap EventBridge envelope
name = body.get("name")
email = body.get("email")Lesson: EventBridge, SQS, SNS, and API Gateway all wrap payloads differently and at different nesting levels. Never assume the shape of the event a Lambda receives — log json.dumps(event) to CloudWatch at the start of every new function during development and inspect the raw structure before writing any field access logic against it.
During teardown, terraform destroy stalled after removing several resources and stopped progressing. The terminal showed Terraform waiting on two resources simultaneously — the SQS queue deletion and the Lambda event source mapping — neither of which completed. After several minutes both operations failed with dependency errors and the destroy exited with a non-zero status, leaving orphaned resources.
The cause was Terraform attempting to delete the SQS queue before the Lambda event source mapping referencing it had been removed. AWS blocks SQS queue deletion while an active event source mapping still points to it. Because the Terraform configuration had no explicit dependency ordering between these two resources — they were linked logically but not via a depends_on reference — Terraform attempted parallel deletion, hit the AWS dependency constraint, and hung.
Fix: Added depends_on = [aws_lambda_event_source_mapping.sqs_trigger] to the aws_sqs_queue resource in main.tf, instructing Terraform to remove the event source mapping before attempting to delete the queue. All 19 resources then destroyed cleanly in a single run.
Lesson: terraform destroy is harder than terraform apply. Terraform builds its dependency graph from explicit resource references in the configuration — if two resources don't reference each other, Terraform has no way to infer their deletion order and will attempt parallel removal. When a destroy hangs, identify which resources are stalled and check whether AWS enforces an implicit deletion dependency between them. Add depends_on to express that dependency explicitly in Terraform.
| Scenario | What happens | Recovery |
|---|---|---|
| lead-handler errors | CloudWatch alarm fires if errors > 1/min | SNS alert sent to operator |
| lead-processor errors | Message stays in SQS — 4-day retention | Automatic retry — zero lead loss |
| DynamoDB write fails | Lambda errors, CloudWatch logs failure | Re-invoke via SQS retry |
| SNS publish fails | Lead stored but no email sent | CloudWatch alarm triggers |
| EventBridge rule disabled | Async path silent — no SQS messages | Re-enable rule via Terraform |
| SQS queue policy missing | EventBridge messages silently dropped | Add aws_sqs_queue_policy resource |
| SNS subscription unconfirmed | Notifications silently discarded | Confirm subscription via email link |
| API Gateway unreachable | Lead source receives 503 | AWS-managed multi-AZ by default |
| Resource | Monthly cost | Notes |
|---|---|---|
| Lambda (both functions) | ~£0.00 | 1M requests/month free tier |
| API Gateway | ~£0.00 | 1M calls/month free tier |
| DynamoDB on-demand | ~£0.03 | 300k requests/month at 10 leads/day |
| SNS | ~£0.00 | 1M publishes/month free tier |
| SQS | ~£0.00 | 1M requests/month free tier |
| EventBridge | ~£0.00 | 14M invocations/month free tier |
| CloudWatch Logs | ~£0.50 | Log storage and metrics |
| Total | ~£0.50/month | vs £150+/month for Salesforce Essentials |
At 100 leads/day: ~£2–3/month. At 1,000 leads/day: ~£8–12/month. Cost scales with usage — zero cost when idle.
This project demonstrated that two processing paths are not redundancy — they serve fundamentally different guarantees. The synchronous path answers the question: how quickly can a lead be stored and an agent notified? The asynchronous path answers a different question: what happens if processing fails? Designing both into a single system — and understanding why each exists — reflects the real tradeoffs in production event-driven architecture.
Silent failures are the hardest class of bug to diagnose. Both the missing IAM permissions and the missing SQS resource policy produced zero visible errors in the AWS console. The only symptoms were absence — an empty queue, an uninvoked function, empty DynamoDB records. Developing the habit of testing each integration boundary explicitly, rather than assuming a correct-looking console status means the system is working, is what made these failures diagnosable rather than mysterious.
Payload structure must be verified at every service boundary. EventBridge, SQS, SNS, and API Gateway each wrap the original payload differently before delivering it to Lambda. Logging the raw event to CloudWatch at the start of every new function during development is not optional — it is the only reliable way to know the exact shape of what the function receives before writing business logic against it.
Decoupling has a concrete cost and a concrete benefit. EventBridge adds a layer of indirection between lead-handler and lead-processor. The cost is additional configuration, an extra resource policy, and one more place for a misconfiguration to hide. The benefit is that a new consumer — an analytics pipeline, a CRM integration, an audit log — can be added by writing a new EventBridge rule without touching handler.py at all. That is architectural extensibility, not theoretical.
Infrastructure that cannot be destroyed is not production-ready. The terraform destroy hang forced an explicit understanding of how AWS enforces deletion dependencies between SQS and Lambda event source mappings. The fix — a single depends_on — is trivial. Finding it required understanding the dependency that AWS enforces but Terraform cannot infer automatically. Every resource has a teardown path, and that path must be as deliberate as the provisioning path.
- API Gateway authorizer — API key or IAM auth on POST /lead to prevent open endpoint abuse
- Input validation — validate and sanitise all fields in lead-handler before storage; reject malformed payloads at the Lambda layer with a 400 response
- Dead-letter queue (DLQ) — capture messages exceeding max SQS retries for manual review and reprocessing; prevents silent lead loss on persistent processor failures
- Separate IAM roles per Lambda — lead-handler needs only
events:PutEvents,dynamodb:PutItem,sns:Publish; lead-processor needs onlysqs:ReceiveMessage,sqs:DeleteMessage,dynamodb:PutItem,sns:Publish - SSM Parameter Store — SNS topic ARN and DynamoDB table name sourced from environment variables at deploy time rather than hardcoded in function code
- lead_id deduplication — pass the handler-generated
lead_idthrough the EventBridge eventdetailso lead-processor reuses the same UUID, preventing duplicate DynamoDB records on retry - GitHub Actions CI/CD — automatic Lambda zip, test, and deploy on push to main, replacing the current manual workflow
- Multi-region — DynamoDB Global Tables for disaster recovery across regions
Sergiu Gota AWS Certified Solutions Architect – Associate · AWS Cloud Practitioner
Built as part of a cloud portfolio to demonstrate production-style serverless event-driven architecture on AWS. Feel free to fork, adapt, or reach out with questions.



















