Sunetra Majumdar
BT + IBM · AI-AUGMENTED SDLC · 2025

Project Stark: Scaling AI-Augmented Business Analysis Across the SDLC

How BT and IBM co-designed an AI-powered Software Development Lifecycle experience — moving from a single proof-of-concept to a production-grade, human-in-the-loop platform for Business Analysts.

120
Stories Generated
45
Working Sessions
15
Stories AI-Drafted / Wk
2
Org Partners: BT + IBM
The Overview

From proof-of-concept to production-grade platform

Project Stark is a joint BT and IBM initiative to embed AI agents throughout the Business Analyst SDLC — from initial idea capture through to engineering-ready user stories and acceptance criteria in JIRA. The programme began as a PoC built on IBM's internal tooling platform, delivered on BT's AWS infrastructure.

AI Agents Human-in-the-Loop SDLC JIRA Integration IBM Design Thinking
Cross-functional team collaborating on AI product strategy
The Challenge

A long, manual, multi-stage process — with compounding friction at every step

Discovery workshops with BT BAs, PMs and the IBM delivery team surfaced a consistent set of pain points across the full idea-to-JIRA workflow.

“Before the process moves — sometimes shaping documents do not include this detail, that it is very important.”

— Senior Business Analyst, BT

Quality

Drafts not “engineering ready,” creating slow, repeated back-and-forth between BAs and Solution Architects.

Inconsistent Output

Shaping documents vary widely in completeness, tone, and complexity, making consistent AI support difficult.

Idle Time

Work frequently stalls in review queues, with no proactive nudge to unblock it.

Time-Consuming Groundwork

Manual AWS process mapping, competitor research, and BA drafting consume disproportionate BA time.

Fragmented Context

Business rules, legal/regulatory constraints (Ofcom, data-sensitive setting) and prior context are scattered rather than centrally available.

Discovery Approach

Design Thinking Workshops — Double Diamond

The team ran a structured series of Design Thinking Workshops (Parts 1–4) following a Double Diamond approach — Discover, Define, Develop, Deliver.

01

As-Is Process Mapping

Mapped the existing BA workflow end to end — idea intake via JIRA, story creation and SA review. Persona, time taken, tools, pain points, and candidate Key Agent Value Areas.

02

Context & Gold-Standard Documents

Defined what “good” looks like for AI inputs and outputs; gold-standard requirement documents, user stories and acceptance criteria scored across quality, topic, and complexity.

03

Value Hypotheses & Prioritisation

Brainstormed and voted on candidate use cases — from AI-checked coverage of story generation to INVEST-aligned drafting, staffing, and QA notification agents.

Sticky notes on a wall during workshop mapping Person writing structured notes on paper Team prioritising ideas in a workshop setting
Solution Design

The AI-Powered SDLC Application

Built around three core views: a personalised dashboard, a human-in-the-loop story review page, and an embedded conversational AI assistant.

View 01

Personalised Work List Dashboard

Each BA gets an immediate view of their queue — stories created, pending reviews, and open flagged issues — alongside a prioritised list surfaced from JIRA tickets.

  • Low-friction triage via Quick-Preview links
  • Priority flags (High/Medium) keep urgent work visible
  • Dedicated SME override-panel for auditable governance
Personalised Work List Dashboard Screen
Personalised Work List Dashboard screenshot
View 02

User Story Review & Feedback Page

The core human-in-the-loop screen. Each AI-generated story is shown with description, acceptance criteria, and a structured quality evaluation reflecting INVEST compliance.

  • Star-rated quality scoring tied to INVEST compliance
  • Given-When-Then acceptance criteria structure
  • Full audit trail from first draft to approved output
 
User Story Review and Feedback Page screenshot
View 03

Embedded AI Assistant

A persistent Smart Assistant panel offers conversational support — answering questions about task status, dependencies, and how to perform common actions.

Embedded AI Assistant Screen
Embedded AI Assistant screenshot

Contextual

Directly aware of the BA's queue and work items.

Conversational

Natural language for status, dependencies, and actions.

Governed

Scoped to authorised actions and BA/PM rules.

Show Cases

Every interaction logged for audit and improvement.

Evaluation Framework

Keeping AI trustworthy — a multi-stage approach

Recognising that AI quality is the biggest adoption risk, the team designed a staged evaluation flow rather than a single pass/fail gate.

Stage 1 · Input evaluation
Guardrail checks (right cluster, right document type) before generation begins.
Agent
Stage 2 · Automated evaluation
Common eval (business intent alignment, traceability, clarity, domain correctness) plus task-specific INVEST criteria.
LLM-as-judge
Stage 3 · End-of-generation review
Qualitative + quantitative feedback scored against a 0–100% threshold; output routed for re-run if below target.
LLM-as-judge
Stage 4 · Human-in-the-loop review
BT and/or IBM human review of high-stakes or low-confidence output before publish.
BA / SME

Test the Limits

Ran the agent close to real-world shaping documents, not just clean demos.

Balanced Judgement

Neither too tough nor too lenient — checked out LLM-as-judge scoring being too generous.

Full Auditability

Score every intermediate step, for auditability and trend-based improvement over releases.

“Project Stark shows what's possible when AI augmentation is designed around trust, not just speed — a human-in-the-loop platform Business Analysts actually want to use.”