All engineering work
Data / CASE STUDY

FinStream

Modern Batch + Streaming Data Platform

A data-engineering platform design bringing event streaming, distributed processing, orchestration and data-quality controls into a shared analytical model.
  • Kafka
  • Spark
  • Airflow
  • dbt
  • PostgreSQL
  • Docker
Engineering projectPublic demo not publishedRepository link not published

Pipeline design is documented in my CV. Detailed operational patterns are design notes; public deployment and performance benchmarks are unverified.

Overview

FinStream brings together batch and streaming pipeline designs using Kafka, Spark, Airflow and dbt, with automated data-quality tests and dimensional models for transaction analytics.

These notes explain the architecture and its trade-offs. Detailed operating guarantees and deployment behavior have not been independently verified.

Problem

Batch reports and streaming applications can disagree when they interpret the same records differently. Late data, duplicate events and undocumented transformations compound the problem. A data platform needs explicit rules for freshness, reconciliation and recovery.

System Architecture

Proposed batch and streaming paths converge on versioned, tested analytical models.SYSTEM DATA FLOW
  1. 01Source systemsEvents · Batch files
  2. 02Event ingestionKafka
  3. 03Data processingSpark
  4. 04Data platformPostgreSQL
  5. 05Quality & modelingdbt
  6. 06AnalyticsTrusted datasets
ADDITIONAL ROUTES
  • Source systemsEvents · Batch filesOrchestrationAirflowBatch
  • OrchestrationAirflowData platformPostgreSQL
  • OrchestrationAirflowQuality & modelingdbt

Kafka captures live events and Spark handles stream processing. Airflow coordinates bounded batch work and backfills. Both paths feed PostgreSQL staging tables, with dbt responsible for versioned transformations, tests and analytical models.

PostgreSQL is a deliberate reference-design choice for an inspectable first implementation. It is not a claim that one database fits every scale or workload.

Engineering Decisions

Define the data contract before tuning infrastructure. A useful contract includes identity, event time, expected freshness, ownership and a clear response to invalid records.

  • Make backfills bounded and repeatable, with explicit processing intervals.
  • Keep business transformations in versioned models rather than duplicating them across consumers.
  • Test uniqueness, required fields and relationships at defined dataset boundaries.
  • Separate pipeline completion from data readiness; a successful job can still publish poor data.

Technology

Kafka and Spark form the streaming path. Airflow coordinates batch dependencies; dbt makes SQL transformations and tests reviewable. PostgreSQL provides a familiar staging and analytics target, with Docker defining repeatable service environments.

  • Kafka
  • Spark
  • Airflow
  • dbt
  • PostgreSQL
  • Docker

Data Flow

Source events enter Kafka and are normalized into staging records. Scheduled extracts use the same record contracts. dbt transforms accepted records into dimensional models and runs quality checks before datasets are marked ready for consumption.

Challenges

The central challenges are aligning event time with processing time, handling schema changes and ensuring that a backfill does not silently double-count records. Reconciliation needs a defined comparison window and source totals, rather than an assumption that two paths eventually agree.

Trade-offs

Spark introduces distributed-processing overhead that smaller workloads may not need. Combining batch and streaming demonstrates useful recovery patterns but increases the surface for operational failure. The implementation should justify each component with workload evidence as it develops.

Observability

Proposed indicators include source-to-table freshness, Kafka lag, processing failures, rejected-row counts and reconciliation differences. Dataset-level quality and lineage should complement infrastructure health so that a green service dashboard does not conceal stale data.

Security

Access should be scoped separately for ingestion, transformation and consumption. Credentials belong outside images and source control. Sensitive attributes should be minimized before analytical publication, with retention and audit requirements defined for each dataset.

Deployment

A containerized reference deployment would allow each service to be inspected independently. Persistent data, Spark checkpoints and orchestration metadata require explicit storage and recovery plans. Production sizing and deployment verification remain future work.

Results

The documented scope covers batch and streaming pipeline design, automated data-quality tests and transaction-analysis models. No throughput, scale or production reliability results are claimed.

A complete validation suite should include duplicate delivery, late events, schema evolution, partial failures and a backfill compared with a known reference dataset.

What I Learned

The design highlights that data quality is an operational property, not just a set of SQL tests. Freshness, reproducibility and reconciliation must be designed together. Further findings should be added after the reference implementation is exercised.

START A CONVERSATION

Let’s build
intelligent systems.

I’m interested in engineering opportunities and collaborations involving production AI, data platforms, backend systems and applied machine learning.