StoryForge MCP Server
Enables AI clients to create and manage Jira tickets for Agile backlog items through natural language commands, integrating structured backlog generation with Jira board actions.
README
StoryForge — AI Backlog & Requirements Assistant
Turn a plain-English business brief into a structured, Agile-ready product backlog (Epics → User Stories → Acceptance Criteria) using an LLM — then measure it, ground it in real documents, and act on it with an agent.
Built as a hands-on tour of the modern Gen-AI product stack: LLM APIs, structured output, prompt engineering, evals, RAG, agents/tools, and MCP.
System map
flowchart LR
brief([Plain-English brief]) --> gen
docs([Company BRD / PRD]) --> rag
subgraph THREE [3 · Ground]
rag[Retrieve relevant chunks]
end
subgraph ONE [1 · Structure]
gen[Generate backlog]
end
rag -. context .-> gen
gen --> backlog[(Structured backlog)]
subgraph TWO [2 · Measure]
eval[Score quality]
end
backlog --> eval
subgraph FOUR [4 · Act]
agent[Agent + tools] --> jira[(Jira board)]
end
backlog --> agent
agent -. exposed via .-> mcp{{MCP server}}
mcp -. any AI client .-> ext([Claude Desktop])
A brief and the company's documents flow in; the app structures a backlog, measures its quality, grounds it in real documents, and acts on it with an agent — reusable by any AI through MCP.
What it demonstrates
Each phase adds one core Gen-AI concept, and each was built to be understood by
running it. Full write-ups (with plain-English explanations) live in
CONCEPTS.md.
| Phase | Concept | What it does |
|---|---|---|
| 1 | Structured output, system prompts, output validation | Brief → typed Backlog object (Epics → Stories → Acceptance Criteria), forced into a Pydantic schema |
| 2 | Evals, golden datasets, regression testing | Rule-based scorers grade backlog quality against a fixed dataset, so prompt changes can be measured, not guessed |
| 3 | RAG, embeddings, vector search, grounding | Retrieves relevant chunks of the company's own documents to ground generation — with a relevance guardrail that refuses out-of-scope questions |
| 4 | Tool calling, agent loops, MCP | An agent pushes backlog stories to a (local) Jira board via tool calling, exposed over MCP for any AI client to use |
Architecture
The app never talks to a specific LLM vendor directly. It talks to an
LLMProvider interface (backend/llm/base.py), and a
factory picks the concrete provider from the LLM_PROVIDER env var. Today that's
Gemini; adding Claude or OpenAI means writing one new class — nothing else
changes. The same seam carries generation, embeddings, and agent tool-calling.
backend/
core/ Phase 1 — schema (models.py), prompt (story_generator.py), CLI
evals/ Phase 2 — golden dataset, scorers, runner
rag/ Phase 3 — chunking, vector store, retriever, Q&A guardrail
agent/ Phase 4 — tools (local Jira), agent loop, MCP server
llm/ provider interface + Gemini implementation
Setup
# 1. Create and activate a virtual environment
python -m venv venv
venv\Scripts\activate # Windows PowerShell
# source venv/bin/activate # macOS/Linux
# 2. Install dependencies
python -m pip install -r requirements.txt
# 3. Add your API key
copy .env.example .env # then edit .env and paste your Gemini key
Get a free Gemini API key at https://aistudio.google.com/apikey
GEMINI_MODEL accepts a comma-separated fallback chain — the app tries each
model in order and fails over if one is overloaded (see .env.example).
Run it, phase by phase
# Phase 1 — brief → structured backlog
python -m backend.core.cli
python -m backend.core.cli "Build an app for tracking gym workouts"
# Phase 2 — evals
python -m backend.evals.test_scorers # offline: scorers vs a bad backlog (no API)
python -m backend.evals.runner myrun # live: generate + score the golden dataset
# Phase 3 — RAG
python -m backend.rag.demo # index a doc → retrieve → grounded backlog
python -m backend.rag.ask "How long before my appointment can I cancel?"
# Phase 4 — agent + tools
python -m backend.agent.tools.jira_demo # the tool alone (no AI)
python -m backend.agent.demo # the agent drives the tool from an instruction
Connect the MCP server to Claude Desktop
Wrap the Jira tools as an MCP server so any MCP client can use them. Add this to
your claude_desktop_config.json, then fully restart Claude Desktop:
{
"mcpServers": {
"storyforge-jira": {
"command": "C:\\path\\to\\storyforge\\venv\\Scripts\\python.exe",
"args": ["-m", "backend.agent.mcp_server"],
"cwd": "C:\\path\\to\\storyforge",
"env": { "PYTHONPATH": "C:\\path\\to\\storyforge" }
}
}
}
Then ask Claude to "create a Jira ticket for X" and watch it land in
jira_board.json.
Roadmap
- [x] Phase 1 — Brief → structured backlog (LLM API, system prompts, structured output)
- [x] Phase 2 — Prompt engineering + evals (measure & improve story quality)
- [x] Phase 3 — RAG (ground stories in uploaded BRD/PRD documents) + guardrails
- [x] Phase 4 — Agents + tools + MCP (push to Jira, agentic refinement)
- [ ] Phase 5 — Web UI + multi-user (the demo shell over everything)
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
mcp-server-qdrant
这个仓库展示了如何为向量搜索引擎 Qdrant 创建一个 MCP (Managed Control Plane) 服务器的示例。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器