Senior Data Engineer · Builder

Renato Aragón

I architect and run cloud data platforms — and ship production software end to end. Senior Data Engineer working with large-scale pipelines on AWS, and founder of Aragón Tecnologia, where I build and operate AI-driven SaaS products serving real customers.

6+shipped products
AWSdata platforms
0→prodfull lifecycle
Renato Aragón
const renato = {
  role: "Senior Data Engineer",
  cloud: ["AWS", "Terraform"],
  builds: "AI SaaS @ Aragón",
  status: "shipping" ✦
}

// about

Data engineering meets product.

I build reliable batch and streaming pipelines, model data for analytics, and automate infrastructure as code. I care about systems that are correct, observable, and maintainable — not just demos that work once.

Through Aragón Tecnologia I take products from idea to production: architecture, development, deployment, and operations — including zero-downtime deploys and AI integrated where it genuinely improves the product.

// stack

The toolkit.

Data & Pipelines

  • Python
  • PySpark
  • Pandas
  • Athena
  • Databricks
  • ETL / ELT
  • Data modelling

Cloud & DevOps

  • AWS
  • Terraform
  • Docker
  • Linux
  • Nginx
  • CI/CD
  • Blue-green deploys

Web & Product

  • Next.js
  • React
  • TypeScript
  • Node.js
  • Tailwind

Storage & AI

  • PostgreSQL
  • Prisma
  • Redis
  • REST APIs
  • LLM integration

// experience

Where I've worked.

2026 — present

Senior Data Engineer

Enterprise Insurance Client, UK (via consultancy)

Large-scale data pipelines and platform engineering on cloud infrastructure — building reliable, observable data flows for enterprise analytics.

ongoing

Founder & Engineer

Aragón Tecnologia

Building and operating a portfolio of AI-driven vertical SaaS products end to end — architecture, development, deployment and operations.

earlier

Data Engineer

BMW Group - USA

Data engineering on AWS — S3, EC2, Athena, PySpark and infrastructure as code with Terraform, in an international, English-speaking environment.

// work

Products I've built & operate.

Vertical SaaS through Aragón Tecnologia — software and data architecture with applied AI, built for trust, security, and precision.

Full portfolio at aragontecnologia.com →

// open source

Engineering in the open.

Public data engineering projects: real patterns, tested and CI-checked, built in small reviewed pull requests. Synthetic data only.

kafka-retail-streaming

Real-time retail events through Kafka (KRaft) and Spark Structured Streaming: windowed aggregations with watermarks, a dead-letter queue, an Iceberg sink with exactly-once semantics, and CDC with Postgres and Debezium.

KafkaSparkCDC

spark-retail-etl

Production-shaped batch ETL with PySpark: incremental loads over a high-watermark, partitioned outputs, quality gates with volume anomaly detection, run summaries for observability, and architecture decision records.

PySparkData QualityCI

dbt-duckdb-analytics

Analytics engineering with dbt and DuckDB: staging to marts with four layers of data tests, from column contracts and referential integrity to cross-layer revenue reconciliation. Runs end to end at zero cost.

dbtDuckDBTesting

terraform-aws-datalake

Reusable Terraform module for an AWS data lake on S3, Glue and Athena: secure by default, plan-time input validation that mirrors the real API constraints, and fmt, validate and tflint gates in CI.

TerraformAWSIaC

data-mapping-framework

Maps heterogeneous sources (CSV, JSON, SQL, Parquet) into a canonical model and runs business rules over it, splitting valid rows from rejects with every violated rule reported per row.

PythonPandasIntegrations

nl-to-sql

Natural language to SQL with Claude: the model proposes, a deterministic guard disposes. A minimal lexer normalizes each query and blocks anything that is not a single read-only SELECT.

LLMDuckDBSafety

meucomboio-pt

A real product in the open: next-train departures and trip planning for Portugal on CP's open GTFS data. Daily sync pipeline with a quality gate, service-calendar aware planning, pricing from CP's public API, and a dbt analytics layer.

ProductGTFSdbt

inventario-familiar

Self-hosted portal for families in an estate process: shared documents, automatic income splitting and individual statements. A dbt + DuckDB analytics layer sits on top, with a reconciliation gate that fails CI if a single cent goes unaccounted for. Synthetic data only.

ProductdbtPostgres

More on github.com/renatoaragon →

// contact

Let's build something.

Open to interesting data engineering and product conversations.