Data Engineering & Cloud Architecture Services

I help teams build data platforms, cloud infrastructure and AI integrations, from architecture and requirements through to production operation. Below you'll find five common data pipeline patterns explained in simple terms, how an engagement with me usually works, and the workshops I offer. My resume lists the projects in detail.

What a data pipeline can look like

Five common patterns, in simple terms. The tools named are examples; the right choice depends on your systems and your team.

Event-driven: process files as they arrive

When a file lands in cloud storage, a notification puts a message on a queue. A small serverless function picks it up, validates and transforms the data and stores the result. There is no server to run, and it scales with the number of files.

  1. Upload file lands in S3
  2. Queue SQS message
  3. Process AWS Lambda: validate · transform
  4. Store S3 · OpenSearch index

Where I've built this: Automotive: validation, ingestion and indexing of vehicle data, with millions of documents indexed in OpenSearch.

Batch: from your systems to a dashboard

Data from databases, APIs and files is collected on a schedule and kept as a raw copy in a data lake. Spark jobs clean and combine it into tables that analysts query with SQL and dashboards read from. Airflow runs every step in the right order and alerts when something fails.

  1. Sources databases · APIs · files
  2. Data lake S3 · Glue Data Catalog · Lake Formation
  3. Transform Spark on AWS Glue or EMR
  4. Query Athena (SQL)
  5. Dashboards reports · BI

Airflow (e.g. Amazon MWAA) schedules and monitors every step · infrastructure as code with Terraform

Where I've built this: Market research: an on-premises data warehouse moved to AWS, with 100+ ETL jobs migrated from Cloudera, a ~1.5 TB data lake and Airflow replacing Oozie.

Real-time: react within seconds

Events such as a click in an app or a reading from a device are written to a stream, processed on the fly and pushed to a fast store or another system. They can be used right away instead of the next day.

  1. Events apps · devices
  2. Stream Kinesis · Kafka
  3. Process Spark Streaming · Kafka Streams
  4. Serve Redis cache · live dashboard · marketing tool

Where I've built this: Entertainment technology: a real-time layer with Spark Streaming, integrated with the batch layer. Media: a real-time Braze integration with under 1 second latency.

AI document extraction: from PDFs to clean data

Documents such as PDFs and scans hold valuable data in unstructured form. The pipeline reads each document, lets an AI model extract the relevant fields, checks the result against rules and existing data, and writes clean records to the database.

  1. Documents PDFs · scans
  2. Extract LLM: Document AI · Vertex AI
  3. Validate rules · existing data
  4. Store structured tables
  5. Use applications · quality dashboard

Where I've built this: Supply chain management: LLM-based document extraction on Google Cloud, processing ~120k documents.

Ask your data in plain English

Instead of writing SQL, people ask a question. An AI assistant such as Claude turns it into a query through an MCP server (Model Context Protocol), which allows read-only access only, checks every query before it runs and limits the load. The answer comes back the same way.

  1. Question plain English
  2. AI assistant Claude
  3. MCP server read-only SQL · EXPLAIN check · rate limiting
  4. Data platform dbt · PostgreSQL · DuckDB

Live example: Shibui Finance makes 60+ years of US stock market data queryable in plain English from Claude and is listed in the official MCP Registry. Rate limiting uses my open-source Caddy module caddy-jsonrpc-matcher.

How an engagement usually works

Typical engagement: I design the platform, build it with your team and hand it over.

1. Design

Together we clarify objectives and requirements, look at the data sources and evaluate the options, including cost and technical feasibility. The result is an architecture with a timeline and milestones. Key decisions are documented as architecture decision records (ADRs), so they can be reviewed and traced later.

2. Build with your team

I set up the cloud infrastructure as code, the data pipelines and the CI/CD pipelines, with monitoring and alerting for production. I work hands-on inside your team and can lead it technically where needed.

3. Hand over

Documentation, training and knowledge transfer along the way mean your team takes over a platform it understands. At the end of the project you get all deliverables agreed upon, including:

  • A repository containing all code and assets required to deploy and run the software
  • A fully automated CI/CD pipeline that provisions the infrastructure, tests and deploys the code
  • Documentation covering usage, maintenance and extension of the deliverables

Workshops

Sometimes the best start is getting your team up to speed. I run workshops tailored to your team's level, with hands-on exercises, for example on:

  • Cloud data platforms: data lake, data warehouse and lakehouse architectures on AWS, Google Cloud or Kubernetes
  • Orchestrating data pipelines with Airflow, including Amazon MWAA
  • Real-time data processing with Kafka, Kafka Streams and Spark Streaming
  • Data management and governance, including GDPR-compliant data residency
  • AI integration: LLM-based document extraction and plain-English data access with MCP

Past workshops include a three-day introduction to the big data ecosystem with hands-on exercises on AWS (healthcare) and a workshop on data management and Lambda architecture (entertainment technology).

Next Steps

Planning a data platform, a pipeline or an AI integration and looking for support? Call me at +49 30 4193 6978, email hello@crichter.io or book a 30-minute call below.

The first conversation is free: we talk through your needs and possible solutions, with no obligation to book me. Take whatever is useful into your own work.

Book meeting