Skip to content

About

Serverless lead routing system for estate agencies — AWS Lambda · API Gateway · SQS · SNS · DynamoDB · EventBridge · Terraform · Python

Resources

Stars

3 stars

Watchers

0 watching

Forks

Latest commit

 

History

30 Commits

Folders and files

Repository files navigation

AWS Property Lead Router

API Gateway · Lambda · DynamoDB · SNS · SQS · EventBridge · Terraform · Python

AWS Lambda Terraform Python DynamoDB EventBridge


Overview

This project builds a fully serverless, event-driven lead routing system on AWS for approximately £0/month at small agency scale. A lead arrives via HTTP POST. Within 3 seconds it is stored in DynamoDB, the right agent is notified by email, and the event is durably queued for asynchronous processing. No human intervention. No inbox. No delay.

The system implements two parallel processing paths intentionally: a synchronous path guarantees immediate storage and notification before the HTTP response is returned, while an asynchronous path through EventBridge and SQS guarantees durability — if the processor fails, the message is retained for up to 4 days and retried automatically. No lead is ever lost.

All 19 AWS resources are provisioned by Terraform in under 2 minutes and destroyed completely with terraform destroy. AWS resources are deployed in the eu-central-1 region.


The Problem

I spent two years at Foxtons, one of London's largest estate agencies. Every day I watched the same thing happen: a lead came in from Rightmove, a website form, or Facebook Ads — and it sat in someone's inbox for hours before an agent picked it up. By the time they called back, the buyer had already spoken to three other agencies.

The fix exists. Salesforce Essentials costs £25 per user per month. For a 6-person agency that's £150/month minimum, plus setup costs, plus training. Most small agencies don't buy it. So leads keep getting lost.

This project builds that fix on AWS for approximately £0/month at small agency scale — fully automated, fully serverless, zero servers to manage.

A lead arrives. Within 3 seconds it is stored, the right agent is notified, and the pipeline records the event. No human intervention. No inbox. No delay.


Why Two Processing Paths?

The system intentionally processes each lead via two parallel paths:

Synchronous path (lead-handler → DynamoDB + SNS): guarantees immediate storage and instant agent notification before the HTTP response is returned. Sub-100ms.

Asynchronous path (EventBridge → SQS → lead-processor): guarantees durability. If lead-processor fails, the message remains in SQS for up to 4 days and retries automatically. No lead is ever lost.

This pattern trades a small amount of duplication for both low latency and high reliability — a deliberate design decision, not an oversight.

Stack: Python 3.11 · boto3 · Terraform · AWS Lambda · API Gateway · DynamoDB · SNS · SQS · EventBridge · CloudWatch · IAM

Region: eu-central-1 · Resources deployed: 19 · Clean destroy: confirmed


Architecture

Architecture Diagram

Lead Source (HTTP POST)
      │
      ▼
API Gateway  ── POST /lead endpoint
      │
      ▼
Lambda: lead-handler  ── Python 3.11 · 29ms avg execution
      │
      ├──▶ DynamoDB (leads-table)  ── Immediate storage · UUID partition key
      │
      ├──▶ SNS (lead-notifications)  ── Instant agent email notification
      │
      └──▶ EventBridge (lead-bus)  ── Async event emission
                  │
                  ▼
            EventBridge Rule (lead-rule)
                  │
                  ▼
            SQS (lead-queue)  ── 4-day retention · SSE-SQS encrypted
                  │
                  ▼
            Lambda: lead-processor  ── SQS consumer · batch size 1
                  │
                  ├──▶ DynamoDB (leads-table)  ── Durable write confirmation
                  │
                  └──▶ SNS (lead-notifications)  ── Pipeline confirmation email

CloudWatch  ── Logs + Metrics + lambda-errors-alarm
IAM         ── Shared execution role · least-privilege per service

Services Used

Service Role
Amazon API Gateway Exposes the public HTTP POST /lead endpoint
AWS Lambda (lead-handler) Ingests lead, writes to DynamoDB, publishes to SNS, fires EventBridge event
AWS Lambda (lead-processor) Consumes from SQS, writes to DynamoDB, publishes pipeline confirmation via SNS
Amazon DynamoDB Stores all lead records — on-demand capacity, UUID partition key
Amazon SNS Delivers real-time email notifications to confirmed agent subscribers
Amazon SQS Durable message buffer between EventBridge and lead-processor — 4-day retention
Amazon EventBridge Custom event bus decoupling lead-handler from downstream consumers
Amazon CloudWatch Logs, metrics dashboard, and lambda-errors-alarm across both functions
AWS IAM Least-privilege execution role granting only required permissions per service
Terraform Provisions and destroys all 19 resources as code

Processing Flow

  1. A lead is submitted via HTTP POST to the API Gateway endpoint with name, email, and message fields
  2. API Gateway proxies the request to the lead-handler Lambda function
  3. lead-handler generates a UUID lead_id, writes the record to DynamoDB, and publishes an SNS notification — all before returning the HTTP 200 response
  4. lead-handler fires a LeadCreated event to the custom EventBridge bus with the lead data in the event detail
  5. The EventBridge rule matches the event on source: lead.api and routes it to SQS
  6. SQS durably queues the message — if lead-processor is unavailable, the message is retained for up to 4 days and retried automatically
  7. lead-processor is triggered by the SQS event source mapping, unwraps the EventBridge envelope from the SQS record body, and writes a confirmation record to DynamoDB
  8. lead-processor publishes a pipeline confirmation notification via SNS
  9. CloudWatch collects invocation metrics, duration, and error counts from both Lambda functions
  10. The lambda-errors-alarm triggers if errors exceed 1 per minute across either function

Project Structure

aws-property-lead-router/
│
├── architecture/
│   └── 00-aws-property-lead-router-architecture.png    # End-to-end architecture overview
│
├── lambda/
│   ├── handler.py                                       # lead-handler: ingest → DynamoDB + SNS + EventBridge
│   └── processor.py                                     # lead-processor: SQS consumer → DynamoDB + SNS
│
├── terraform/
│   ├── main.tf                                          # All 19 AWS resource definitions
│   └── outputs.tf                                       # API Gateway endpoint URL
│
├── screenshots/
│   ├── 01-project-structure.png
│   ├── 02-terraform-init.png
│   ├── 03-dynamodb-table-active.png
│   ├── 04-sns-topic-created.png
│   ├── 05-sns-subscription-confirmed.png
│   ├── 06-lambda-handler-deployed.png
│   ├── 07-api-gateway-route.png
│   ├── 08-api-gateway-terraform-apply.png
│   ├── 09-api-test-curl-success.png
│   ├── 10-dynamodb-leads-stored.png
│   ├── 11-email-notification-received.png
│   ├── 12-cloudwatch-handler-logs.png
│   ├── 13-sqs-queue-created.png
│   ├── 14-lambda-processor-deployed.png
│   ├── 15-sqs-trigger-enabled.png
│   ├── 16-eventbridge-rule-active.png
│   ├── 17-dynamodb-full-pipeline-6-items.png
│   ├── 18-email-sqs-pipeline-confirmed.png
│   ├── 19-cloudwatch-metrics-dashboard.png
│   ├── 20-cloudwatch-alarm-configured.png
│   └── 21-terraform-destroy-complete.png
│
├── .gitignore
└── README.md

Deployment

Prerequisites

  • AWS CLI configured with appropriate credentials
  • Terraform >= 1.0 installed
  • Docker (optional — for local Lambda testing)

Provision

terraform init        # Download hashicorp/aws v5.100.0 provider
terraform plan        # Preview all 19 resources before apply
terraform apply       # Provision full stack (~2 minutes)

After apply, confirm the SNS email subscription by clicking the link in the AWS confirmation email sent to your agent address. The stack is not fully operational until the subscription is confirmed.

Test

curl -X POST https://<your-api-id>.execute-api.eu-central-1.amazonaws.com/lead \
  -H "Content-Type: application/json" \
  -d '{"name":"Test Lead","email":"test@email.com","message":"Interested in a 2-bed flat"}'

Expected response: HTTP 200 with a confirmation message. Within 3 seconds, the lead appears in DynamoDB and an email notification arrives at the confirmed subscriber address.

Teardown

terraform destroy     # Remove all 19 resources · zero idle cost after teardown

Screenshots

1. Terraform Init

hashicorp/aws v5.100.0 installed and locked. Provider lock file committed for reproducible deployments.

Terraform Init


2. DynamoDB Table Active

leads-table provisioned via Terraform. Partition key: lead_id (UUID). Capacity mode: on-demand. Status: Active.

DynamoDB Table


3. SNS Topic Created

lead-notifications topic with confirmed email subscription. Agent notification pipeline active.

SNS Topic


4. SNS Subscription Confirmed

Email subscription confirmed via AWS confirmation link — notifications now delivered to the agent inbox.

SNS Confirmed


5. Lambda Handler Deployed

lead-handler deployed — Python 3.11. Writes to DynamoDB, publishes to SNS, fires to EventBridge in a single invocation.

Lambda Handler


6. API Gateway Route

POST /lead route configured on lead-api with Lambda integration wired.

API Gateway Route


7. Live API Test

curl -X POST with real lead data returning 200 OK and confirmation message.

API Test


8. DynamoDB Leads Stored

4 lead records in leads-table after initial testing — all with UUID partition keys and ISO 8601 timestamps.

DynamoDB Items


9. Email Notification Received

SNS → Gmail confirmed. Real agent notifications delivered within seconds of the API call.

Email


10. CloudWatch Handler Logs

lead-handler execution — Duration: 29.49ms, Billed: 479ms, Memory: 128MB used: 87MB. Zero errors.

CloudWatch Logs


11. SQS Queue Created

lead-queue — Standard type, SSE-SQS encryption, 4-day message retention.

SQS Queue


12. Lambda Processor Deployed

lead-processor deployed — Python 3.11. SQS event source mapping configured. Consumes from lead-queue, writes to DynamoDB, publishes to SNS.

Lambda Processor


13. SQS Trigger Enabled

SQS event source mapping active: lead-queue → lead-processor. State: Enabled. Batch size: 1.

SQS Trigger


14. EventBridge Rule Active

lead-rule on lead-bus — Status: Enabled. Routes events from source: lead.api into lead-queue.

EventBridge


15. DynamoDB Full Pipeline — 6 Items

6 items after full pipeline test — including entries confirming the EventBridge → SQS → lead-processor → DynamoDB path works end-to-end.

DynamoDB 6 Items


16. Email After SQS Pipeline

Pipeline confirmation email triggered by the full async path — EventBridge → SQS → lead-processor → SNS → Gmail working.

Email SQS


17. CloudWatch Metrics Dashboard

Lambda metrics — Invocations, Duration, Errors, Throttles across both functions in a single view.

CloudWatch Metrics


18. CloudWatch Alarm Configured

lambda-errors-alarm — Errors > 1 for 1 datapoint within 1 minute. Status: Insufficient data (correct — no errors fired).

CloudWatch Alarm


19. Terraform Destroy — Clean Teardown

terraform destroy — all 19 resources destroyed. Zero orphaned infrastructure. Zero idle cost.

Terraform Destroy


Engineering Decisions

Why Lambda over EC2 or Fargate? Lead routing is event-driven and spiky — a form is submitted, a function runs for 29ms, then it is idle. Lambda is the correct compute model. Fargate would be correct for a persistent web application. EC2 for sustained predictable load. The access pattern here is pure event-driven: Lambda wins on cost, simplicity, and zero operational overhead.

Why DynamoDB over RDS? The access pattern is simple: write a lead record, read it by lead_id. No joins. No relational queries. DynamoDB on-demand at this volume costs cents per month. RDS has a fixed monthly baseline cost — the wrong tool for a key-value access pattern.

Why EventBridge instead of direct Lambda-to-SQS? EventBridge decouples the producer (lead-handler) from downstream consumers. The handler emits an event without knowing what processes it. Rules route events to SQS, which acts as a durable buffer. This allows adding new consumers later — analytics, CRM integrations, audit pipelines — without modifying handler.py. Direct Lambda-to-SQS would tightly couple ingestion to a single consumer.

Why SQS between EventBridge and lead-processor? SQS provides durability that direct Lambda invocation cannot. If lead-processor errors, the message stays in the queue and retries automatically for up to 4 days. Zero leads lost even if the processor is down for hours. This is the reliability layer of the system.

Why a custom EventBridge bus? Using lead-bus instead of the default AWS event bus isolates application events from AWS service events in the same account. Cleaner observability, easier to scale, and future event patterns are less likely to conflict with AWS-generated events.

Why Terraform over console provisioning? Every resource is defined in code. terraform apply provisions the full stack in under 2 minutes. terraform destroy removes all 19 resources completely. No manual steps, no configuration drift, no orphaned resources. The destroy step is as important as the apply — if infrastructure cannot be cleanly removed, it is not production-ready.

Why a shared IAM role? Both Lambda functions share a single execution role for simplicity in this project. In production, separate roles per function would enforce strict least-privilege. Documented in Production Improvements.

Why no VPC, IGW, NAT, ALB, or subnets? All services used are AWS-managed and require no VPC placement. Lambda has native outbound internet access. Adding a VPC would cost ~£30/month for a NAT Gateway and introduce latency with zero security benefit for this architecture. The absence of VPC is a deliberate design decision, not an omission.


Troubleshooting

1. Lambda Not Receiving SQS Messages Due to Missing IAM Permissions

After deploying the SQS event source mapping and confirming it was enabled in the console, lead-processor was not being invoked when messages arrived in the queue. The SQS messages were visible in the console and the queue depth was growing, but Lambda was making no attempt to consume them. No errors appeared in the Lambda console — the function was simply silent.

The root cause was a missing IAM permission. The execution role for lead-processor did not include sqs:ReceiveMessage or sqs:DeleteMessage. Lambda requires both permissions to poll a queue and acknowledge processed messages. Without them, the event source mapping is enabled but functionally inert — Lambda cannot pull from the queue and produces no visible error at the function level. The failure was only visible in CloudWatch Logs as an access denied exception on the polling attempt.

Fix: Added explicit sqs:ReceiveMessage, sqs:DeleteMessage, and sqs:GetQueueAttributes to the IAM role policy in main.tf and re-ran terraform apply. The event source mapping began delivering messages to lead-processor immediately after the policy update.

Lesson: IAM permissions must match the exact API calls a service makes — not just the service name. Lambda polling SQS requires three distinct permissions. Always verify the CloudWatch Logs for the function itself when an event source mapping appears enabled but the function is not triggering — the access denied exception is logged there, not in the Lambda console.


2. EventBridge Messages Silently Dropped — No Delivery to SQS

With the EventBridge rule active and confirmed enabled, messages published to lead-bus were not arriving in lead-queue. The SQS queue remained empty after every API call, the lead-processor was never invoked, and no error appeared anywhere in the console. The EventBridge rule showed as enabled with the correct target configured.

The cause was a missing SQS resource-based policy. EventBridge requires explicit sqs:SendMessage permission granted on the target queue via a queue policy — this is separate from the Lambda execution role and is not inherited from any other IAM configuration. Without it, EventBridge attempts delivery and silently discards every message. This was the hardest failure to diagnose in the entire project because no error was surfaced anywhere in the EventBridge or SQS consoles.

Fix: Added an aws_sqs_queue_policy Terraform resource explicitly granting EventBridge sqs:SendMessage on lead-queue, scoped to the EventBridge service principal and the specific rule ARN. Messages routed correctly to the queue on the next apply.

Lesson: EventBridge → SQS always requires an explicit queue resource policy. It is not inherited from any IAM role. Silent failure with no visible error is the only symptom. Test the full event routing path end-to-end immediately after provisioning by sending a real event and verifying the SQS queue depth increases — do not assume the configuration is correct based on console status alone.


3. lead-processor Storing Empty Records Due to EventBridge Envelope Wrapping

After the EventBridge → SQS → Lambda path was working end-to-end, lead records were appearing in DynamoDB from lead-processor but all fields were empty — name, email, and message were None for every record written by the async path. The synchronous path via lead-handler was writing correctly; the issue was isolated to lead-processor.

The cause was an incorrect assumption about the payload structure. When EventBridge routes an event to SQS, the original JSON body is wrapped inside an EventBridge envelope. The SQS record body is not the raw lead data — it is the full EventBridge event object with the lead data nested under a detail key. The initial processor.py was parsing record['body'] and accessing name, email, and message directly, which returned None for all fields because they don't exist at the top level of the envelope.

Fix: Updated processor.py to unwrap the envelope before accessing lead fields:

event_body = json.loads(record['body'])
body = event_body["detail"]  # unwrap EventBridge envelope
name = body.get("name")
email = body.get("email")

Lesson: EventBridge, SQS, SNS, and API Gateway all wrap payloads differently and at different nesting levels. Never assume the shape of the event a Lambda receives — log json.dumps(event) to CloudWatch at the start of every new function during development and inspect the raw structure before writing any field access logic against it.


4. Terraform Destroy Hanging on SQS and Lambda Event Source Mapping

During teardown, terraform destroy stalled after removing several resources and stopped progressing. The terminal showed Terraform waiting on two resources simultaneously — the SQS queue deletion and the Lambda event source mapping — neither of which completed. After several minutes both operations failed with dependency errors and the destroy exited with a non-zero status, leaving orphaned resources.

The cause was Terraform attempting to delete the SQS queue before the Lambda event source mapping referencing it had been removed. AWS blocks SQS queue deletion while an active event source mapping still points to it. Because the Terraform configuration had no explicit dependency ordering between these two resources — they were linked logically but not via a depends_on reference — Terraform attempted parallel deletion, hit the AWS dependency constraint, and hung.

Fix: Added depends_on = [aws_lambda_event_source_mapping.sqs_trigger] to the aws_sqs_queue resource in main.tf, instructing Terraform to remove the event source mapping before attempting to delete the queue. All 19 resources then destroyed cleanly in a single run.

Lesson: terraform destroy is harder than terraform apply. Terraform builds its dependency graph from explicit resource references in the configuration — if two resources don't reference each other, Terraform has no way to infer their deletion order and will attempt parallel removal. When a destroy hangs, identify which resources are stalled and check whether AWS enforces an implicit deletion dependency between them. Add depends_on to express that dependency explicitly in Terraform.


Failure Scenarios and System Behaviour

Scenario What happens Recovery
lead-handler errors CloudWatch alarm fires if errors > 1/min SNS alert sent to operator
lead-processor errors Message stays in SQS — 4-day retention Automatic retry — zero lead loss
DynamoDB write fails Lambda errors, CloudWatch logs failure Re-invoke via SQS retry
SNS publish fails Lead stored but no email sent CloudWatch alarm triggers
EventBridge rule disabled Async path silent — no SQS messages Re-enable rule via Terraform
SQS queue policy missing EventBridge messages silently dropped Add aws_sqs_queue_policy resource
SNS subscription unconfirmed Notifications silently discarded Confirm subscription via email link
API Gateway unreachable Lead source receives 503 AWS-managed multi-AZ by default

Cost Model

Resource Monthly cost Notes
Lambda (both functions) ~£0.00 1M requests/month free tier
API Gateway ~£0.00 1M calls/month free tier
DynamoDB on-demand ~£0.03 300k requests/month at 10 leads/day
SNS ~£0.00 1M publishes/month free tier
SQS ~£0.00 1M requests/month free tier
EventBridge ~£0.00 14M invocations/month free tier
CloudWatch Logs ~£0.50 Log storage and metrics
Total ~£0.50/month vs £150+/month for Salesforce Essentials

At 100 leads/day: ~£2–3/month. At 1,000 leads/day: ~£8–12/month. Cost scales with usage — zero cost when idle.


What I Learned

This project demonstrated that two processing paths are not redundancy — they serve fundamentally different guarantees. The synchronous path answers the question: how quickly can a lead be stored and an agent notified? The asynchronous path answers a different question: what happens if processing fails? Designing both into a single system — and understanding why each exists — reflects the real tradeoffs in production event-driven architecture.

Silent failures are the hardest class of bug to diagnose. Both the missing IAM permissions and the missing SQS resource policy produced zero visible errors in the AWS console. The only symptoms were absence — an empty queue, an uninvoked function, empty DynamoDB records. Developing the habit of testing each integration boundary explicitly, rather than assuming a correct-looking console status means the system is working, is what made these failures diagnosable rather than mysterious.

Payload structure must be verified at every service boundary. EventBridge, SQS, SNS, and API Gateway each wrap the original payload differently before delivering it to Lambda. Logging the raw event to CloudWatch at the start of every new function during development is not optional — it is the only reliable way to know the exact shape of what the function receives before writing business logic against it.

Decoupling has a concrete cost and a concrete benefit. EventBridge adds a layer of indirection between lead-handler and lead-processor. The cost is additional configuration, an extra resource policy, and one more place for a misconfiguration to hide. The benefit is that a new consumer — an analytics pipeline, a CRM integration, an audit log — can be added by writing a new EventBridge rule without touching handler.py at all. That is architectural extensibility, not theoretical.

Infrastructure that cannot be destroyed is not production-ready. The terraform destroy hang forced an explicit understanding of how AWS enforces deletion dependencies between SQS and Lambda event source mappings. The fix — a single depends_on — is trivial. Finding it required understanding the dependency that AWS enforces but Terraform cannot infer automatically. Every resource has a teardown path, and that path must be as deliberate as the provisioning path.


Production Improvements

  • API Gateway authorizer — API key or IAM auth on POST /lead to prevent open endpoint abuse
  • Input validation — validate and sanitise all fields in lead-handler before storage; reject malformed payloads at the Lambda layer with a 400 response
  • Dead-letter queue (DLQ) — capture messages exceeding max SQS retries for manual review and reprocessing; prevents silent lead loss on persistent processor failures
  • Separate IAM roles per Lambda — lead-handler needs only events:PutEvents, dynamodb:PutItem, sns:Publish; lead-processor needs only sqs:ReceiveMessage, sqs:DeleteMessage, dynamodb:PutItem, sns:Publish
  • SSM Parameter Store — SNS topic ARN and DynamoDB table name sourced from environment variables at deploy time rather than hardcoded in function code
  • lead_id deduplication — pass the handler-generated lead_id through the EventBridge event detail so lead-processor reuses the same UUID, preventing duplicate DynamoDB records on retry
  • GitHub Actions CI/CD — automatic Lambda zip, test, and deploy on push to main, replacing the current manual workflow
  • Multi-region — DynamoDB Global Tables for disaster recovery across regions

Author

Sergiu Gota AWS Certified Solutions Architect – Associate · AWS Cloud Practitioner

GitHub LinkedIn

Built as part of a cloud portfolio to demonstrate production-style serverless event-driven architecture on AWS. Feel free to fork, adapt, or reach out with questions.

About

Serverless lead routing system for estate agencies — AWS Lambda · API Gateway · SQS · SNS · DynamoDB · EventBridge · Terraform · Python

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages