Data Pipeline & ETL Engineer

data-engineer

Builds Redis task queues, background workers, ETL processes, and analytics stores.

Where it installs

# .claude/agents/data-engineer.md
---
name: data-engineer
description: ...
tools: Read, Write, Edit, Glob, Grep, Skill, Bash, TodoWrite
model: inherit
---

What it does

Data Pipeline & ETL Engineer is a Cursor agent playbook named data-engineer. AgenticKit installs it so a session can adopt this role instead of acting as a generic coding assistant.

The data-engineer agent handles background task execution and high-throughput data processing. It sets up robust queues with retry backoff, dead-letter queues, CSV/JSON batch ingestion, and analytical database sync.

Builds Redis task queues, background workers, ETL processes, and analytics stores. Data engineer for Cursor that builds asynchronous job queues (BullMQ/Redis), ETL data pipelines, batch processors, and analytics schemas.

Why it exists

A pile of agent files is not a team. data-engineer exists so one job, data pipeline & etl engineer, has a written brief, required skills, and a stop condition.

It ships in the Engineering Kit. Skills it must read: Asynchronous Background Jobs & Task Queues; PostgreSQL & Drizzle ORM Schema Standards; SaaS Architecture Patterns & Multi-Tenancy. Commands that adopt this role: /data-pipeline: Background Job & Task Queue Architecture; /build-api: API Endpoint & Route Generator; /sre-incident: SRE Incident Responder & Root Cause Analysis.

When to use it

  • Offloading heavy computations to background workers
  • Importing and processing large user data files
  • Synchronizing operational data to analytics stores

When not to use it

  • Do not ask data-engineer to do a different role's job. If you need a launch post, switch agents.
  • Do not keep the same chat after this agent has finished its artifact. Start a reviewer in a new thread.
  • Do not invoke every agent in the kit for a small change.

Example workflow

  1. Install the Engineering Kit so .cursor/agents/data-engineer is on disk.
  2. Run /data-pipeline: Background Job & Task Queue Architecture, or start a chat and tell Cursor to adopt the data-engineer role.
  3. The agent should read: Asynchronous Background Jobs & Task Queues; PostgreSQL & Drizzle ORM Schema Standards; SaaS Architecture Patterns & Multi-Tenancy.
  4. It produces the artifact for this role only, then stops.
  5. A different agent or you review. Same-chat self-review is not a review.

Example usage

Example: "You are data-engineer. Offloading heavy computations to background workers. Read data-pipelines before you edit."

Or let the pipeline invoke it: /data-pipeline: Background Job & Task Queue Architecture.

Capabilities you should actually see: Asynchronous job queue design with BullMQ and Redis

Example output

  • data-engineer should leave files or a written verdict, not a vibe check. Asynchronous job queue design with BullMQ and Redis High-volume data transformation and ETL pipeline scripts
  • Engineering agents should touch the slice they were given (schema, route, test, or review note) and nothing else.

Best practices

  • Asynchronous job queue design with BullMQ and Redis
  • High-volume data transformation and ETL pipeline scripts
  • Batch CSV/JSON file ingestion and streaming parsers
  • Dead-letter queue handling and automated job retries
  • One role per chat unless a command is explicitly orchestrating a sequence.

Common mistakes

  • Using data-engineer as a synonym for 'the Cursor agent'. It is a brief, not the product.
  • Skipping the required skills and hoping the role name is enough.
  • Letting the writer approve its own PR.
  • Invoking this agent and three unrelated ones in the same prompt.

Frequently asked questions

  • What is the data-engineer Cursor agent?
    Data Pipeline & ETL Engineer: Data engineer for Cursor that builds asynchronous job queues (BullMQ/Redis), ETL data pipelines, batch processors, and analytics schemas.
  • When should I invoke data-engineer?
    Offloading heavy computations to background workers Importing and processing large user data files Synchronizing operational data to analytics stores
  • What skills does data-engineer use?
    Asynchronous Background Jobs & Task Queues; PostgreSQL & Drizzle ORM Schema Standards; SaaS Architecture Patterns & Multi-Tenancy
  • How do I run data-engineer in Cursor?
    Run /data-pipeline: Background Job & Task Queue Architecture or /build-api: API Endpoint & Route Generator or /sre-incident: SRE Incident Responder & Root Cause Analysis, or start a chat and adopt the data-engineer role.
  • Is data-engineer the same as Cursor's built-in Agent?
    No. Cursor Agent is the product harness. This file is a specialist brief you install so that harness takes a named role.

Install data-engineer with the rest of the team

46 agents, 61 skills, and 47 slash commands, installed into .cursor/ and .claude/. Engineering and marketing kits, one license.