LLM Integration & Streaming AI Workflows

.cursor/skills/llm-integration

Anthropic, OpenAI, streaming tokens, structured JSON, and fallback routing.

Where it installs

# .claude/skills/llm-integration/SKILL.md
---
name: llm-integration
description: ...
---
# identical content, skills are the same file in both editors

What it does

LLM Integration & Streaming AI Workflows is a Cursor skill: a SKILL.md playbook that AgenticKit installs at .cursor/skills/llm-integration. Cursor's agent loads it when the task matches the skill description, instead of stuffing the same rules into every chat.

Best practices for integrating Large Language Models into web apps, including streaming UI, prompt templating, token budgeting, and fallback providers.

Anthropic, OpenAI, streaming tokens, structured JSON, and fallback routing. The file is domain knowledge, not a slash macro. You do not type it. The agent reads it when the job is llm integration.

Why it exists

Default Cursor has no memory of how your team ships code, schema, and tests. Founders paste the same conventions into chat, then watch the model drift on the next turn. This skill exists so those conventions live on disk and load only when relevant.

It ships in the Engineering Kit. It is wired to the AI & Vector Engineer (ai-engineer) and Senior Backend Architect (backend-architect) agents. It is wired to the /ai-feature: AI Integration & Vector Pipeline Builder and /llm-chat: Streaming AI Chat Assistant Interface commands. That graph is the point: the skill is the standard, the agent is the role, the command is the trigger.

When to use it

  • Use LLM Integration & Streaming AI Workflows when the work is specifically about anthropic, openai, streaming tokens, structured json, and fallback routing.
  • Load it before a long session that will touch this surface more than once. A one-line edit does not need the full playbook.
  • Use it on production SaaS work (multi-tenant apps, paid features, anything that will be reviewed) rather than a throwaway prototype.

When not to use it

  • Do not treat the skill as a replacement for a slash command. Skills teach. Commands run a pipeline.
  • Do not paste the whole SKILL.md into chat. If Cursor is not loading it, fix the description or invoke the matching command.
  • Do not use it as a generic 'write better code' rule. Scope is this domain only.

Example workflow

  1. Install the Engineering Kit so .cursor/skills/llm-integration lands in the repo.
  2. Open a new Cursor agent chat pointed at the files this skill governs.
  3. Ask the AI & Vector Engineer (ai-engineer) to read the skill, or run /ai-feature: AI Integration & Vector Pipeline Builder.
  4. Review the first artifact against the practices below. If it violates one, stop and correct the file. Do not prompt 'just finish it'.
  5. Commit the skill-guided files with the rest of the slice so the next session inherits the same standard.

Example usage

Example prompt: "Read .cursor/skills/llm-integration and apply it to this change. Do not invent extra conventions."

Or trigger the pipeline: /ai-feature: AI Integration & Vector Pipeline Builder. That command is written to load this skill.

Check the output against: Stream AI responses using Vercel AI SDK to minimize perceived latency

Example output

  • The skill itself does not write a single output file. It changes what the agent is allowed to produce in code, schema, and tests.
  • When you run /ai-feature: AI Integration & Vector Pipeline Builder, expect repo files plus notes under docs/, not a chat-only answer.

Best practices

  • Stream AI responses using Vercel AI SDK to minimize perceived latency
  • Enforce structured JSON output using OpenAI response_format or Anthropic tool use
  • Implement token usage counters and rate limit throttling per user tier
  • Add retry logic with exponential backoff for rate-limited API calls

Common mistakes

  • Ignoring "Stream AI responses using Vercel AI SDK to minimize perceived latency" and hoping a later prompt will clean it up.
  • Copying the skill into .cursor/rules as always-on. That burns context and fights Cursor's load-on-match design.
  • Running two overlapping skills that contradict each other in the same turn.
  • Letting the agent skip tests or types because 'the skill is about architecture'.

Frequently asked questions

  • What is the LLM Integration & Streaming AI Workflows Cursor skill?
    Best practices for integrating Large Language Models into web apps, including streaming UI, prompt templating, token budgeting, and fallback providers. It lives at .cursor/skills/llm-integration after you install the Engineering Kit.
  • When should I use llm-integration instead of a Cursor rule?
    Use a rule for always-on or glob-scoped constraints. Use this skill for the full playbook that should load only when the task matches.
  • Which agents read llm-integration?
    AI & Vector Engineer (ai-engineer); Senior Backend Architect (backend-architect). Those roles are told to open this file before they edit.
  • Which commands use llm-integration?
    /ai-feature: AI Integration & Vector Pipeline Builder; /llm-chat: Streaming AI Chat Assistant Interface.
  • Does this skill work outside AgenticKit?
    Yes. A SKILL.md in .cursor/skills/llm-integration is a normal Cursor skill. AgenticKit is the packaged version plus the agent and command graph.

Add llm-integration and the other 48 Cursor skills

AgenticKit installs 61 skills, 46 agents, and 47 slash commands into .cursor/. One license, lifetime updates.