New Cohort Starts:

Donate
← Back to the skill map
Data & Storage/Data & services

Data Ingestion

5 micro-topics in Data & Storage, each with the evidence that proves you have it and the written reason behind every link. Before you start, it rests on 4 other domains. Downstream, it holds up 3 domains.

5 topics  ·  depth 7–10 of 16  ·  4 internal links  ·  4 in  ·  3 out

Before you start

Load-bearing links first. Each one says what in this domain rests on what outside it, and why.

Data Modeling

1 link

load-bearing

Build a repeatable ingestion pipeline rests on Choose primary and foreign keys

Idempotent loads require a key to upsert on.

Python Core

1 link

load-bearing

Ingest from files and APIs rests on Read and write files

Ingestion reads files.

Production Python

1 link

load-bearing

Observability

1 link

supporting

Schedule and monitor batch jobs rests on Configure alerts

Unattended jobs need alerting to be trustworthy.

What you will be able to do

In prerequisite order. Each idea names the artifact that closes it, the market demand that put it on the map, and what it stands on.

Depth 7  ·  procedural  ·  unassisted  ·  M22 §22.1

Ingest from files and APIs

Pulls data from CSV, JSON and HTTP sources

Evidence
Loads an external dataset into a table
Market anchor
Data engineering
Rests on
load-bearing

Read and write files · Python Core

Ingestion reads files.

load-bearing

Call HTTP APIs asynchronously · Production Python

Ingestion pulls from APIs.

Depth 8  ·  procedural  ·  unassisted  ·  M22 §22.1

Clean and normalize records

Handles nulls, duplicates and inconsistent types

Evidence
Turns a messy export into a loadable dataset
Market anchor
Data engineering
Rests on
load-bearing

Ingest from files and APIs

You clean what you have loaded.

Depth 9  ·  conceptual  ·  scaffolded  ·  M22 §22.1

Build a repeatable ingestion pipeline

Makes ingestion idempotent and resumable

Evidence
Reruns a failed load without duplicating rows
Market anchor
Data engineering · ETL
Rests on
load-bearing

Clean and normalize records

A pipeline formalizes ingestion and cleaning.

load-bearing

Choose primary and foreign keys · Data Modeling

Idempotent loads require a key to upsert on.

Depth 10  ·  conceptual  ·  scaffolded  ·  M22 §22.1

Checkpoint long-running jobs

Records progress so work resumes

Evidence
Resumes a job after a crash mid-batch
Market anchor
Data engineering
Rests on
load-bearing

Build a repeatable ingestion pipeline

Checkpointing makes a pipeline resumable.

Depth 10  ·  meta  ·  guided  ·  M22 §22.1

Schedule and monitor batch jobs

Runs work on a cadence with failure alerts

Evidence
A failed nightly job pages someone
Market anchor
Data engineering · operations
Rests on
load-bearing

Build a repeatable ingestion pipeline

Scheduling automates a pipeline that already runs.

supporting

Configure alerts · Observability

Unattended jobs need alerting to be trustworthy.

What rests on this domain

Everything downstream that names an idea here as a prerequisite, grouped by where it lives.

AI Safety & Governance

1 link

load-bearing

Recognize data poisoning risk rests on Build a repeatable ingestion pipeline

Poisoning targets the ingestion path.

Embeddings & Vectors

1 link

load-bearing

Chunk documents for retrieval rests on Clean and normalize records

You chunk cleaned documents.

Agents

1 link

supporting

Checkpoint and resume agent state rests on Checkpoint long-running jobs

Resumability is the same pattern as in batch jobs.

Retool. Retrain. Relaunch.

375 ideas. 17 weeks. No tuition, ever.

Vets Who Code is a veteran-run 501(c)(3). The accelerator is free, remote, and we don’t take a share of your first paycheck.

Free · Remote · 17 weeks · EIN 86-2122804