Business Context: Why Performance Reviews Fail and How AI Fixes the Right Parts
The fundamental problem with performance reviews is not that they happen too often or too rarely. It is that they are too time-consuming for managers to do well, and too generic in output to be genuinely useful for employee development. A manager responsible for 8 direct reports who must write substantive quarterly reviews is spending 16-24 hours per quarter on review writing alone. When that time is constrained (and it always is), quality suffers. Reviews become formulaic, focus on the most recent events (recency bias), and fail to synthesise patterns across the full period. 360 feedback collection is even more manual: aggregating qualitative responses from 5-8 peers per employee, identifying themes, and presenting them in a usable format requires significant administrative work that usually falls on HR. AI changes what is feasible in this process. It does not improve performance. It reduces the administrative burden of documenting and communicating about performance, so managers can spend their review time on the conversation rather than the writing. Under the EU AI Act, AI systems used in employment evaluation contexts may be classified as high-risk under Annex III, which means documentation of the system's purpose, limitations, and human oversight requirements is expected. SpeedMVPs builds this documentation into the delivery.
Architecture: Feedback Collection, Theme Analysis, and Narrative Generation
The system has three functional layers. The collection layer provides a lightweight feedback form for peers to complete: a rating scale per competency dimension and a free-text box for qualitative observations. Forms are distributed via email with a unique link per respondent-subject pair, and responses are anonymous to the subject (visible only to the manager and HR). Supabase handles authentication, data storage, and Row Level Security to enforce access controls on sensitive feedback data. The analysis layer processes the collected feedback: quantitative ratings are aggregated per competency, qualitative responses are analysed by Claude to extract themes, specific examples, and development observations. Claude identifies patterns that appear across multiple respondents (which carry more weight) and distinguishes them from observations that appear in only one response (which may reflect individual working relationship dynamics more than actual performance patterns). The narrative generation layer uses Claude to produce three outputs: a synthesised 360 feedback summary for the manager, a manager review draft that weaves together the 360 themes with manager observations, and a development highlights document that focuses on growth areas identified across the feedback cycle. All three outputs are editable by the manager before finalisation.
AI Components: Theme Extraction and Structured Narrative Drafting
Claude handles two distinct language tasks. First, theme extraction from qualitative 360 responses: given 5-8 peer responses for an employee, identify the recurring themes (both strengths and development areas), extract specific examples that illustrate each theme, and assess the consistency of observations across respondents. This is a genuine AI advantage over manual aggregation: Claude can identify that four different respondents are describing the same underlying behaviour in different words, which a manual aggregator working under time pressure would miss. Second, narrative drafting: given the competency ratings, extracted themes, and the manager's own observations (entered via the interface), generate a structured performance review draft. The draft covers: summary of the period's contributions, key strengths with examples, development areas with suggested actions, and overall rating justification. The draft is written in a professional, constructive tone calibrated to the company's performance review style guide, which SpeedMVPs configures during the build.
Challenges: Bias, Anonymity, and Employment Law
AI-assisted performance review raises legitimate concerns about bias and fairness. LLMs can reflect the biases present in their training data, which may include patterns that disadvantage certain groups. SpeedMVPs builds in two mitigations. First, the AI is applied identically across all employees in the review cycle, without access to demographic data that would allow differential treatment. Second, the AI output is reviewed and approved by the manager before it is finalised. The EU AI Act's classification of employment evaluation AI as potentially high-risk creates documentation expectations that SpeedMVPs addresses in the delivery package: a description of the system's function, its known limitations, and the human review requirement. Anonymity of 360 feedback respondents must be technically enforced. Even well-intentioned managers can identify respondents from distinctive writing styles or specific examples. The system stores responses with respondent IDs visible only to the HR admin, not to the manager or the subject employee. For UK employers, the Employment Rights Act and Equality Act 2010 apply to how performance information is used in employment decisions. The AI-generated review is evidence supporting a decision, not the decision itself, and managers must be trained accordingly.
Outcomes: Time Savings and Review Quality Improvements
HR teams and managers that use AI-assisted performance review tools report the same primary outcome: the time managers spend writing reviews drops by 60-70%. A review that previously took 45-60 minutes to write from scratch takes 15-20 minutes to review, edit, and finalise from an AI draft. At scale, for a company with 100 employees going through a quarterly review cycle, this represents hundreds of manager-hours returned to productive work per cycle. Secondary outcomes include more consistent review quality (the AI draft ensures all reviews cover the same structural elements, reducing the variance between strong-writer managers and weak-writer managers) and higher-quality 360 feedback summaries (multi-respondent theme extraction surfaces patterns that manual aggregation misses).
Lessons: Manager Training Is as Important as the Technology
AI-generated review drafts are only useful if managers understand how to use them well. The risk is that managers approve AI drafts without meaningfully engaging with them, which undermines the purpose of the review. SpeedMVPs recommends building a brief manager training module into the product: a short guided experience that walks managers through how to evaluate an AI draft, what to personalise, and what to verify against their own observations. The second lesson is to design for the manager's workflow, not the HR team's. HR teams love process; managers tolerate it. The review interface should be as low-friction as possible: clear status, minimal clicks from feedback collection to finalised review, and mobile-accessible for managers who do reviews during commutes or between meetings.