
Katie Kinkel · Available for work
AI drafts. Human QA. Better UX.
AI QA, LLM evaluation, human-in-the-loop review, and AI-assisted operations — available for remote part-time and contract work.
Available for work · Open to remote · Part-time · Contract · AI evaluation / QA / operations
Previous QA experience (2019–2022): Spotify support program via 24-7 Intouch — human QA only; AI work starts in 2025.
AI Evaluation
What Katie evaluates on every model draft — before anything ships.
Instruction following
Did the model do what was asked — and only what was asked?
Factual accuracy
Are claims checkable and true, or does the draft invent details?
Completeness
Are required sections, fields, and next steps present?
Tone
Is the voice appropriate for the audience and brand?
Formatting
Is structure scannable — headings, lists, links, and layout ready to use?
Hallucinations
Flag invented facts, fake citations, and confident wrong answers.
Edge cases
Probe odd inputs, missing context, and failure modes before ship.
Multi-output comparison
Score several model drafts against the same rubric and pick a winner.
Final human review
The non-negotiable gate: Katie signs off before anything leaves.
Hire focus areas
AI Evaluation
Rubrics, side-by-side scoring, and clear pass/fail judgment on model output.
LLM QA
Catch hallucinations, tone misses, incomplete answers, and instruction drift.
AI-Assisted Operations
Turn messy notes and requests into organized trackers, minutes, and briefs.
Human-in-the-Loop Review
Every AI draft gets a human pass. Judgment is the product that ships.
Spotify Support / QA experience
Via 24-7 Intouch (2019–2022): QA audits, alignments, nesting assistance, Guru process feedback, LIO & Marquee ROTAs, and the CrS Phoenix Hub — 3+ years Creative Support / QA Buddy. Human QA only; AI tools came later (2025+).
Case studies
AI projects are from 2025 onward (Problem → AI used → what Katie reviewed → result). Spotify Support / QA (2019–2022) is earlier human QA proof — no AI in that role.
LLM evaluation demo — support email rewrite
- Problem
- Need a clear customer-support email rewrite that follows instructions, stays factual, and sounds on-brand — without inventing policy details.
- AI used
- 2025 portfolio exercise: same prompt to ChatGPT, Claude, and Gemini (not client-confidential).
- What Katie reviewed / corrected
- Scored each draft on instruction following, factual accuracy, completeness, tone, formatting, and hallucinations. Flagged invented policy language and incomplete next steps; improved the prompt and picked a corrected final.
- Result
- Documented rubric scores, failure notes, and a shippable final email after human review.

Rubric scoring three model drafts for a support email rewrite 
Side-by-side comparison of model outputs with failure flags 
Final human-reviewed support email ready to ship AI-assisted operations — messy notes to verified minutes
- Problem
- Scattered meeting notes, action items, and side chats needed a single accurate deliverable volunteers could trust.
- AI used
- 2025: ChatGPT + Cursor to draft structured minutes and an action tracker from raw notes.
- What Katie reviewed / corrected
- Checked names, dates, votes, and owners line by line; removed invented attendees and vague “someone will…” actions; fixed formatting for scannability.
- Result
- Verified minutes and a clear follow-up list — AI for speed, human QA for reliability.

Messy raw notes and request fragments before organization 
Organized minutes and action tracker after human QA Spotify support QA via 24-7 Intouch
3+ years Spotify CS / QA- Problem
- Creative Support needed consistent quality across nesting, process docs, and live work — not just ticket volume.
- AI used
- None — human QA only (2019–2022). No AI tools in this role.
- What Katie reviewed / corrected
- QA audits and alignments, nesting assistance, Guru process feedback, LIO & Marquee ROTAs, and maintenance of the CrS Phoenix Hub team resource.
- Result
- 3+ years in Spotify CS / QA Buddy path via 24-7 Intouch — quality systems and team resources that held up under review.
Legal aid support communications
- Problem
- A disabled person needed accurate paperwork and communications without risky AI inventions.
- AI used
- 2025: ChatGPT for first-pass drafts of letters and forms language.
- What Katie reviewed / corrected
- Line-by-line factual and tone review before anything was sent; corrected incomplete fields and overconfident phrasing.
- Result
- Reviewed communications that stayed accurate and respectful under human sign-off.
Gordonston Neighborhood Association minutes
- Problem
- Association meetings needed reliable minutes on a volunteer timeline.
- AI used
- 2025: AI for speed on first drafts of minutes and agendas.
- What Katie reviewed / corrected
- Human QA for accuracy of motions, attendance, and action owners before publishing.
- Result
- Clear, neighbor-friendly minutes maintained as GNA secretary.
Local cat rescue coordination
- Problem
- Rescue logistics and outreach notes were easy to drop under volunteer load.
- AI used
- 2025: AI-assisted drafting for outreach notes and coordination lists.
- What Katie reviewed / corrected
- Verified details, contacts, and next steps so volunteers could move without missing care items.
- Result
- Faster day-to-day coordination with fewer dropped details.
Website prototypes & proofreading
- Problem
- Stakeholders needed something real to react to — not another vague brief.
- AI used
- 2025: Cursor / Copilot / ChatGPT for site prototypes and copy variants.
- What Katie reviewed / corrected
- Proofreading and UX clarity pass so pages ship clean, readable, and on-message.
- Result
- Prototypes ready for feedback with a human quality gate.
Skills & tools
Skill-first: evaluation, judgment, QA, research, operations, and communication. Surfaces Katie works in daily: Cursor · Copilot · ChatGPT. Model families stay unversioned because tools rotate; QA judgment does not.
- Cursor
- GitHub Copilot
- Claude
- Gemini
- Notion
Spotify Support / QA via 24-7 Intouch
- LLM evaluation
- Human QA & judgment
- HITL review
- Research
- AI operations
- Communication
- Cursor
- ChatGPT
- Claude
Site stack
This portfolio runs on a modern web stack — useful context for AI QA work with product and engineering teams.
- Next.js
- React
- TypeScript
- Tailwind CSS
- Vercel
- Framer Motion
- next-themes
- WebMCP
- Chrome Prompt API
- GitLab
How I work with AI
Brief → draft → QA → UX. AI speeds the first pass; Katie's review is the reliability gate for clarity, accuracy, and usable experience.

“AI generates; Katie evaluates, corrects, organizes, and ships — human judgment is the quality gate.”
What I bring
Curiosity
Learns new AI tools quickly — then directs them with a clear evaluation plan.
Ownership
Owns the final product. AI drafts; Katie reviews, corrects, and ships.
Reliability
AI generates; Katie evaluates accuracy, tone, and completeness before anyone depends on it.
Purpose
Professional and clear — helping teams move faster without losing the human quality gate.
AI Evaluation & QA
The #1 capability: score model output against a rubric, catch failures, and decide what ships.
LLM Workflows
Design brief → draft → review loops in Cursor, Copilot, and ChatGPT that stay human-gated.
AI Operations
Structure chaos into agendas, folders, trackers, and follow-ups teams can actually run.
Research
Pull sources fast, compare options, and verify before anyone acts on an AI summary.
Prototyping
Spin up site mockups and copy variants so stakeholders react to something real.
Communication
Minutes, emails, briefs, and stakeholder language — clear, warm, and on-brand.
Looking for remote AI QA, LLM evaluation, AI training, or AI operations help?
Katie is available for part-time, contract, and project-based work.
Or email directly: ktkinkel@gmail.com
Available for remote part-time and contract work in AI QA, LLM evaluation, and AI-assisted operations — where human review is the product.
- 3+
Years Spotify CS / QA
- 3
AI surfaces in daily use
- 6
Model families in use
- 3+
Community ops roles
© 2026 Katie Kinkel