flowchart

Data Pipeline Flowchart

Map a nightly extract, validate, transform, and load job, including a quarantine branch for bad rows and a dashboard refresh at the end.

Data Pipeline Flowchart — Mermaid flowchart template preview
Static preview — open the editor for a live, editable version.

Open in editor

The code

flowchart TD
  A[(Source database)] --> B[Extract batch]
  B --> C[Validate schema]
  C --> D{Records valid?}
  D -- No --> E[Quarantine bad rows]
  E --> F[Alert data team]
  D -- Yes --> G[Transform fields]
  G --> H[Enrich with reference data]
  H --> I[Load into warehouse]
  I --> J[Refresh dashboards]
  F --> K((Run finished))
  J --> K

How this template works

This template maps a batch data job from source to dashboard: extract from the operational database, validate the schema, transform and enrich the records, load them into the warehouse, and refresh the dashboards that executives read in the morning. Data teams use it to document the nightly job for on-call rotations, and it is equally useful as a design artifact when proposing a new pipeline, because the quarantine branch forces the conversation about bad data before it happens in production.

The syntax introduces the cylinder. The first line, flowchart TD, sets the top-down direction, and the source node is written A[(Source database)] — square brackets wrapped in parentheses render a database cylinder instead of a rectangle, which instantly tells the reader this node is storage rather than an action. Every other node is a plain rectangle for a processing step. The validation gate D{Records valid?} splits the flow: clean records continue down the transform path, while D -- No --> E[Quarantine bad rows] diverts failures into a holding area, and F[Alert data team] makes sure the diversion is visible. Both paths converge on the terminator K((Run finished)), drawn with double parentheses — the run finishes even when rows were quarantined, and the diagram says so honestly.

The gotcha is the cylinder bracket order. The shape is [( on the left and )] on the right; write A[()] or A([Source]) and you get a different shape or a parse error. The second trap is labels with commas: a step like H[Enrich with reference data] is safe, but H[Enrich with region, currency data] must be quoted as H{"Enrich with region, currency data"} or the parser will misread the label. Finally, keep node ids unique — pipelines often have repeated step names, and reusing an id merges two distinct stages into one node.

To adapt it, rename the steps to your actual job stages and add a node for your orchestrator if scheduling matters. If bad rows should block the run rather than skip, point the quarantine branch at a failed terminator instead.

Related templates: the order fulfillment flowchart for the operational process that produces this data, the database transaction sequence for the row-level guarantees behind the extract, and the swimlane process flowchart for assigning pipeline stages to owning teams. The flowchart diagram guide lists every node shape including the cylinder.

Variations to try

Related templates