The generative AI projects worth building in 2026 are small, finished apps that call a language model, ground it in real data and measure whether the answers are right. All 25 ideas below follow that rule, grouped by level. If you only have time for one, choose by what you need it to do for you.
For beginners, the PDF Chat Assistant (project 2). One document, one embedding model and one prompt teach you the whole retrieval loop in a weekend, and you finish with a Gradio app, a 20 question test set and a short note on where it got answers wrong.
For a final-year project, the Multilingual Policy Q&A Assistant (project 14). It tackles a genuinely Indian problem, English documents queried in Hindi or a regional language, which is the kind of grounded work final-year examiners reward. You submit a RAG app with citations, an evaluation report and a project report chapter on limitations.
For a developer portfolio, the MCP Server and Tool-Using Agent (project 19). It shows you can give a model safe access to real systems, a core skill in agent engineering roles. Expect to ship a working MCP server, a client agent, logs and a permission model. It also maps well to the skills in our guide on how to become an AI engineer.
For advanced agentic work, the LLM Evaluation Harness (project 20). Few student portfolios include one, yet production teams rely on them. The output is a test set, scored runs across two models and a CI job that fails when quality drops.
Key Highlights of Generative AI Projects
- 25 builds in three levels: 8 beginner, 9 intermediate RAG projects and 8 advanced agentic AI projects, each with focus, stack, deliverables, skills and what it proves to a recruiter.
- A 2026 tech stack table covering models, embeddings, vector databases, orchestration, MCP, evaluation and deployment, with official documentation links.
- Real cost drivers explained (input tokens, output tokens, embeddings, hosting) so you can estimate a monthly bill before you build.
- A GitHub folder template and README checklist that reviewers can scan in under a minute.
- Evaluation, guardrails and deployment treated as projects in their own right, since that is where most portfolios fall short.
Introduction to Generative AI Projects
Any app where a model creates something (text, code, an image caption, a spoken reply) while your code decides what goes in and checks what comes out counts as a generative AI project. If you'd like a refresher on what generative AI is and how it works, read that first; this article assumes you know the basics and want to build.
Plenty of gen AI projects online stop at "call the API and print the reply". That was fine in 2023. Today, recruiters and project guides expect the model to be grounded in your own data, with some way to measure quality and some thought given to cost and safety, so this list is built around those expectations rather than reading like a generic top 10 artificial intelligence projects roundup.
You'll need working Python for almost all of them. If that's your gap, the Python Programming training covers what you need before project 1. Vector search, agents and the rest you can pick up as you go. The wider family of AI projects includes classic machine learning too, but this list sticks to LLM, RAG and agent builds.
What Makes a Good Generative AI Project for Your Portfolio
It answers a real question, uses data you're allowed to use and comes with evidence that it works. A generic chatbot proves little. One that answers questions about your college's exam rulebook, cites the page and is scored on a test set you wrote yourself proves a lot.
Four qualities separate the strong projects from the weak ones:
- A narrow problem. "Answer questions about the Motor Vehicles Act" beats "legal assistant".
- Grounding. The model works from documents, a database or tools, so its answers can be checked.
- Measurement. A small test set with expected answers, run every time you change a prompt or model.
- A visible trade-off. You tried two chunk sizes, two models or two prompts, and you can explain why you kept one.
Generative AI Tech Stack for Projects in 2026
Every project below draws on some mix of the layers in this table. Start with the beginner column and move to the advanced one only when a real limit forces you to.
| Layer | Beginner choice | Advanced choice | Official reference |
| Model | A hosted LLM API (Claude, GPT, Gemini) or a small open model run locally | Several models behind a router, plus a fine-tuned open model | Ollama quickstart for local models |
| Embeddings | The embedding endpoint from your model provider or a small open embedding model | A domain-tuned or multilingual embedding model | Provider docs |
| Vector store | Chroma, running in-process | PostgreSQL with pgvector and an HNSW index | pgvector on GitHub |
| Orchestration | Plain Python functions | LangGraph for stateful, multi-step agents | LangGraph overview |
| Tool access | Native tool use (function calling) in the model API | Model Context Protocol (MCP) servers | MCP server concepts |
| Evaluation | A CSV of questions and expected answers, scored by hand | Ragas metrics plus an LLM judge in CI | Ragas metrics |
| Interface and deploy | Gradio | FastAPI backend in a Docker container | Framework docs |
Ollama serves a local API on http://localhost:11434 that also accepts OpenAI-style and Anthropic-style requests, so swapping a local model for a hosted one is a matter of changing a base URL. pgvector supports cosine distance with the <=> operator and offers HNSW and IVFFlat indexes; its README notes HNSW supports vectors up to 2,000 dimensions, so check your embedding size before committing to it. LangGraph calls itself a low-level orchestration framework for long-running, stateful agents, with durable execution and human-in-the-loop checkpoints. For a single prompt that's overkill, but for LLM projects with branching steps it fits well.
What will all this cost? Hosted model bills come from input tokens, output tokens and any per-call charges for server-side tools. Anthropic's tool use documentation notes that defining tools adds a tool use system prompt (286 tokens for Claude Opus 5.5) on top of your own tool schemas, so a long tool list costs you on every call. Embedding a corpus is a one-time cost each time the corpus changes. Prices move often, so check the provider's pricing page when you budget, and log token counts from day one rather than guessing.
25 Generative AI Projects: Quick Comparison Table
| # | Project | Level | Core stack | Est. build time | Key output | Key skill |
| 1 | Prompt-Templated Email Drafter | Beginner | LLM API, Jinja templates | 3 to 5 hours | Drafts in three tones | Prompt design |
| 2 | PDF Chat Assistant | Beginner | Embeddings, Chroma, Gradio | 1 to 2 days | Q and A with page citations | Basic RAG |
| 3 | Meeting Notes to Action Items | Beginner | LLM API, structured outputs | 1 day | JSON action list | Schema-bound output |
| 4 | Local Chatbot with an Open Model | Beginner | Ollama, Gradio | 4 to 6 hours | Offline chat app | Running open models |
| 5 | Resume and Job Description Matcher | Beginner | Embeddings, LLM API | 1 day | Match score and gap list | Similarity search |
| 6 | Natural Language to SQL Helper | Beginner | LLM API, SQLite | 1 to 2 days | Safe read-only queries | Tool calling basics |
| 7 | Alt Text Generator for Images | Beginner | Vision-capable LLM | 1 day | Alt text for a folder | Multimodal input |
| 8 | Lecture Notes Quiz Generator | Beginner | LLM API, Pydantic | 1 day | MCQs with answer keys | Output validation |
| 9 | Multi-Document Knowledge Base with Citations | Intermediate | pgvector, FastAPI | 3 to 5 days | Cited answers across 100+ files | Chunking, retrieval |
| 10 | Hybrid Search RAG with Reranking | Intermediate | BM25, vectors, reranker | 3 to 4 days | Before and after retrieval scores | Retrieval tuning |
| 11 | Support Ticket Triage and Reply Drafter | Intermediate | LLM API, classifier prompts | 3 days | Labelled queue, draft replies | Classification with LLMs |
| 12 | Invoice and Receipt Data Extractor | Intermediate | Vision LLM, structured outputs | 3 to 4 days | Validated JSON and error log | Extraction accuracy |
| 13 | LoRA Fine-Tune of a Small Open Model | Intermediate | Transformers, PEFT | 4 to 6 days | Adapter weights and comparison | Fine-tuning trade-offs |
| 14 | Multilingual Policy Q&A Assistant | Intermediate | Multilingual embeddings, RAG | 1 week | Hindi or regional language answers | Cross-lingual retrieval |
| 15 | Lecture and Podcast Summariser | Intermediate | Whisper, LLM API | 3 days | Timestamped summaries | Speech to text pipelines |
| 16 | Codebase Documentation Generator | Intermediate | LLM API, AST parsing | 4 days | Generated docs per module | Long context handling |
| 17 | Contract Clause Comparator | Intermediate | RAG, structured outputs | 4 to 5 days | Clause diff table | Comparative reasoning |
| 18 | Multi-Agent Research Assistant | Advanced | LangGraph, search tool | 1 to 2 weeks | Cited research brief | Agent orchestration |
| 19 | MCP Server and Tool-Using Agent | Advanced | MCP SDK, LLM API | 1 week | Server plus agent with logs | Tool access design |
| 20 | LLM Evaluation Harness | Advanced | Ragas, pytest, CI | 1 week | Scored runs and CI gate | Evaluation |
| 21 | Guardrails and Safety Layer | Advanced | Classifiers, PII redaction | 1 to 2 weeks | Blocked attack log | LLM security |
| 22 | Voice Agent for Appointment Booking | Advanced | Whisper, LLM, TTS | 2 weeks | Spoken booking flow | Latency engineering |
| 23 | Pull Request Code Review Agent | Advanced | GitHub API, LLM API | 1 to 2 weeks | Review comments on PRs | Agent in a workflow |
| 24 | Self-Correcting Agentic RAG | Advanced | LangGraph, retrieval grader | 1 to 2 weeks | Fewer wrong answers | Reflection loops |
| 25 | Cost-Aware LLM Router with Caching | Advanced | FastAPI, cache, two models | 1 week | Cost and latency dashboard | LLM operations |
Build times assume a few focused hours a day and that you've already finished the previous level.
Generative AI Projects for Beginners (8 Projects)
Each of these generative AI projects for beginners uses one model call pattern and a small amount of data. Keep the scope tight and put the time you save into a README and a handful of test cases.
1. Prompt-Templated Email and Message Drafter
You give it three bullet points and pick a tone (formal, friendly or firm), and it returns a polished email. Store the prompts in versioned template files rather than hard-coded strings.
Focus: Separating instructions, examples and user input inside a prompt.
Tech Stack: Python, any hosted LLM API, Jinja2 templates, a simple CLI or Gradio form.
Deliverables:
- Three prompt templates with two few-shot examples each.
- A side-by-side table of outputs for the same input across tones.
- A short note on which instruction changes made the biggest difference.
Skills practised: Prompt structure, few-shot examples, temperature settings.
What it proves to a recruiter: You treat prompts as code that can be versioned and compared.
2. PDF Chat Assistant
Upload one PDF, a syllabus for instance, and ask questions about it. Behind the scenes the app chunks and embeds the text, pulls back the closest four chunks and tells the model to answer only from those, with page numbers.
Focus: The full retrieval loop on a single document.
Tech Stack: pypdf, an embedding model, Chroma, an LLM API, Gradio.
Deliverables:
- A working chat app that shows the retrieved chunks next to each answer.
- A 20 question test set with expected answers and page numbers.
- A result line, for example "16 of 20 correct", plus notes on the four misses.
Skills practised: Chunking, embeddings, similarity search, grounded prompting.
What it proves to a recruiter: You understand why RAG works and how it fails.
3. Meeting Notes to Action Items
Paste in a raw meeting transcript and get back a JSON list of action items, each with an owner, a due date if one was mentioned and the sentence it came from. OpenAI's Structured Outputs guide explains how a supplied JSON Schema keeps responses valid, and other providers offer similar schema-bound modes.
Focus: Forcing a model to return data your code can trust.
Tech Stack: LLM API with a JSON schema, Pydantic, a small Gradio page.
Deliverables:
- A Pydantic model for an action item.
- Ten sample transcripts and their extracted output.
- A count of missed or invented action items against your own reading.
Skills practised: Schema design, source quoting, precision versus recall thinking.
What it proves to a recruiter: You can connect LLM output to downstream systems without breaking them.
4. Local Chatbot with an Open Model
Run an open model on your own machine and put a chat interface on it that works offline. According to the Ollama quickstart, ollama run gemma4:e2b downloads and starts a model; that one needs roughly a 7.2 GB download, and Ollama recommends 8 GB of VRAM or Mac unified memory for it.
Focus: Understanding what runs locally, what it costs in memory, and how quality compares with a hosted model.
Tech Stack: Ollama, its local API on port 11434, Gradio.
Deliverables:
- A chat app with conversation history.
- A comparison of five prompts answered by the local model and one hosted model. Our Claude vs ChatGPT vs Gemini comparison can help you pick the hosted one.
- A table of response time per prompt on your hardware.
Skills practised: Model serving basics, context windows, hardware trade-offs.
What it proves to a recruiter: You can work with open models in settings where data can't leave the building.
5. Resume and Job Description Matcher
Embed a resume and ten job descriptions, rank the jobs by similarity, then have a model list the missing skills for the top three. Use your own resume or a sample one; other people's profiles are off limits.
Focus: Combining vector similarity with a generative explanation.
Tech Stack: An embedding model, NumPy for cosine similarity, an LLM API.
Deliverables:
- A ranked list with similarity scores.
- A gap report per job with specific missing skills.
- A note on cases where the score and the model's judgement disagreed.
Skills practised: Embeddings, cosine similarity, explanation prompts.
What it proves to a recruiter: You know embeddings measure closeness rather than suitability, and you can explain why that matters.
6. Natural Language to SQL Helper
Load a sample SQLite database (a small sales dataset works) and let users ask things like "Which region sold most in March?" The model writes SQL through a tool call, and your code confirms it's a SELECT, runs it over a read-only connection and shows both the query and the result.
Focus: Letting a model act through a tool while your code keeps control.
Tech Stack: SQLite, an LLM API with tool use, sqlparse for checks.
Deliverables:
- A tool definition with a JSON schema for the query.
- A blocklist test showing DROP and UPDATE are rejected.
- 25 test questions with expected results and a pass rate.
Skills practised: Tool calling, input validation, least privilege.
What it proves to a recruiter: You put safety checks between a model and a database.
7. Alt Text Generator for Images
Point a script at a folder of images, generate short descriptive alt text for each one and write the results to a CSV. Plenty of websites skip this accessibility work, so it's a useful thing to have done properly.
Focus: Sending images to a vision-capable model and keeping descriptions short and factual.
Tech Stack: A vision-capable LLM API, Pillow for resizing, a CSV writer.
Deliverables:
- Alt text for 30 images with character limits enforced.
- A manual review column marking each description as accurate, vague or wrong.
- Before and after examples where a prompt change fixed a common error.
Skills practised: Multimodal prompting, image preprocessing, human review.
What it proves to a recruiter: You can handle non-text inputs and judge output quality honestly.
8. Lecture Notes Quiz Generator
Take a chapter of notes and produce ten multiple choice questions, each with one correct answer, three plausible distractors and the source line it came from. Every question gets validated against a Pydantic schema, and any whose answer isn't in the source text is thrown out.
Focus: Generating content and automatically checking it.
Tech Stack: LLM API, Pydantic, simple string matching or a second model call as a checker.
Deliverables:
- A quiz JSON file and a printable version.
- A rejection log showing questions the checker removed.
- Feedback from two classmates on question quality.
Skills practised: Validation, generator and checker pattern, educational content design.
What it proves to a recruiter: You don't trust model output by default, and you built a way to catch its errors.
Intermediate Generative AI and RAG Projects (9 Projects)
Most strong student portfolios sit at this level, where the data grows, retrieval needs tuning and you start measuring everything. If the pattern itself is still hazy, read how retrieval-augmented generation works before you start the RAG projects from project 9 onwards.
9. Multi-Document Knowledge Base with Citations
Index 100 or more documents, public policy PDFs or a product's documentation for example, and answer questions with citations down to file and section. Keeping the vectors in PostgreSQL with pgvector means metadata filters such as "only 2025 circulars" run in the same query as the similarity search.
Focus: Chunking strategy, metadata, and citation accuracy at scale.
Tech Stack: PostgreSQL, pgvector with an HNSW index, an embedding model, FastAPI, an LLM API.
Deliverables:
- An ingestion script that records file, page and section per chunk.
- A comparison of two chunk sizes (for example 300 and 800 tokens) on a 40 question test set.
- An API endpoint that returns answer, citations and retrieval scores.
Skills practised: Vector indexing, SQL plus vector queries, citation design.
What it proves to a recruiter: You can run retrieval on a real database instead of a notebook.
10. Hybrid Search RAG with Reranking
Pure vector search often struggles with exact terms like section numbers. So add keyword search (BM25) alongside the vector search from project 9, merge the two result lists and rerank the top 20 with a cross-encoder.
Focus: Measuring retrieval quality on its own, before generation.
Tech Stack: rank_bm25 or PostgreSQL full-text search, your vector store, an open reranker model.
Deliverables:
- Hit rate at 5 for vector only, keyword only and hybrid with reranking.
- Five example queries where hybrid fixed a miss.
- Latency added by the reranker per query.
Skills practised: Information retrieval metrics, fusion methods, latency budgeting.
What it proves to a recruiter: You debug RAG at the retrieval layer instead of blaming the model.
11. Support Ticket Triage and Reply Drafter
The system sorts support tickets by category and urgency, then drafts a reply grounded in a help centre article. Nothing goes out until a human approves it.
Focus: LLM classification with clear labels and a human approval step.
Tech Stack: LLM API, a small labelled dataset, a simple review queue built in Gradio or a web framework.
Deliverables:
- 100 labelled sample tickets (write them yourself or use a public dataset).
- A confusion matrix for category predictions.
- A review screen with approve, edit and reject actions logged.
Skills practised: Label design, classification metrics, human-in-the-loop workflows.
What it proves to a recruiter: You can automate part of a business process without removing human control.
12. Invoice and Receipt Data Extractor
Pull vendor, date, GSTIN, line items and totals out of photographed receipts and PDF invoices. Then check the totals yourself by summing the line items in code, and flag every mismatch.
Focus: Field-level accuracy and arithmetic checks the model cannot fake.
Tech Stack: A vision-capable LLM with schema-bound output, Pydantic validators, pandas.
Deliverables:
- A labelled set of 30 receipts (your own, with personal data removed).
- Field-level accuracy per field, not one overall number.
- An error log with the image, the wrong value and the correct one.
Skills practised: Document understanding, validation rules, error analysis.
What it proves to a recruiter: You know extraction is only as good as the checks around it.
13. LoRA Fine-Tune of a Small Open Model
Use LoRA to teach a small open model one narrow format, such as writing commit messages from diffs. Hugging Face's PEFT library trains only a small set of extra parameters, and that's what makes the job feasible on a single consumer GPU or a free cloud notebook.
Focus: Knowing when fine-tuning beats prompting, and when it does not.
Tech Stack: Hugging Face Transformers, PEFT, a dataset of 500 to 2,000 examples, a GPU notebook.
Deliverables:
- Training script, config and adapter weights pushed to a model repo.
- A comparison of the base model with a prompt, against the fine-tuned model, on 50 held-out examples.
- A cost and time log for the training run.
Skills practised: Dataset preparation, LoRA settings, held-out evaluation.
What it proves to a recruiter: You can make an evidence-based call on fine-tuning rather than following hype.
14. Multilingual Policy Q&A Assistant
Index English government scheme documents and let people ask questions in Hindi, Tamil or another regional language, then answer in that same language with citations back to the English source. Because the problem is local and real, it makes a strong final-year choice.
Focus: Cross-lingual retrieval and answer faithfulness.
Tech Stack: A multilingual embedding model, pgvector or Chroma, an LLM API with good Indian language support, Gradio.
Deliverables:
- A 30 question test set per language, written or checked by a native speaker.
- Retrieval hit rate when questions are embedded directly versus translated to English first.
- A limitations section listing schemes or phrasings it handles poorly.
Skills practised: Multilingual NLP, translation trade-offs, evaluation with native speakers.
What it proves to a recruiter: You build for real users, including those who don't read English.
15. Lecture and Podcast Summariser
Transcribe a long lecture or podcast with Whisper, then generate a timestamped summary and a list of key terms. OpenAI's Whisper repository lists models from 39M to 1,550M parameters, needing anywhere from about 1 GB to about 10 GB of VRAM, with code and weights released under the MIT licence.
Focus: Chaining speech to text with long-input summarisation.
Tech Stack: Whisper (start with a small model), ffmpeg, an LLM API, a map-reduce summary approach.
Deliverables:
- Transcripts and summaries for three recordings of 45 minutes or more.
- Timestamps that link back to the audio.
- A comparison of transcription quality between two Whisper model sizes on one recording.
Skills practised: Audio pipelines, chunked summarisation, model size trade-offs.
What it proves to a recruiter: You can handle long, messy input and keep traceability.
16. Codebase Documentation Generator
Parse a small open source repository with Python's ast module, then generate docstrings for each function and a module overview for each file. It's the sort of chore developer teams would happily hand off.
Focus: Feeding the right code context to a model without exceeding limits.
Tech Stack: Python ast, an LLM API with a large context window, MkDocs for output.
Deliverables:
- A generated documentation site for one repository.
- A reviewer checklist scoring 20 generated docstrings for accuracy.
- Token usage per file, so you can see how cost scales.
Skills practised: Code parsing, context packing, documentation standards.
What it proves to a recruiter: You can apply LLMs to a developer team's daily work.
17. Contract Clause Comparator
Upload two versions of a rental agreement and get back a table of changed, added and removed clauses with plain-language explanations, along with a note that it isn't legal advice.
Focus: Aligning sections across documents before asking the model to compare.
Tech Stack: Section splitting with regex or headings, embeddings for alignment, an LLM API with structured output.
Deliverables:
- A clause-by-clause diff table in HTML.
- Five test document pairs with known changes and a detection rate.
- A note on clauses it misaligned and why.
Skills practised: Document alignment, comparative prompts, structured reporting.
What it proves to a recruiter: You can design a pipeline where the model does one careful job inside a larger system.
Advanced Agentic AI Projects (8 Projects)
In agentic AI projects you hand the model a goal, some tools and a loop, and it decides the next step itself. Control, observability and cost are the hard parts. Before settling on a framework, compare the orchestration frameworks such as LangGraph and CrewAI so your choice has a reason behind it. These AI agent projects suit developers who already have at least two finished intermediate builds. If the idea is new, read what agentic AI is and how agentic AI differs from generative AI.
18. Multi-Agent Research Assistant
The graph has three roles. A planner breaks a question into sub-questions, a researcher searches and reads sources, and a writer turns the findings into a cited brief, with LangGraph's checkpoints letting you pause before the writer step for human review.
Focus: State management and hand-offs between agents.
Tech Stack: LangGraph, an LLM API, a web search tool or API, a citation checker.
Deliverables:
- A graph diagram and the code for each node.
- Five research briefs with every claim linked to a source.
- A trace log showing each step, tool call and token count.
Skills practised: Agent orchestration, state, human-in-the-loop checkpoints.
What it proves to a recruiter: You can build agents whose every step can be inspected.
19. MCP Server and Tool-Using Agent
Write an MCP server that exposes a small system, say a library catalogue or your own task tracker, and connect an agent to it. The MCP documentation defines three server building blocks: tools the model can call, resources that provide read-only context, and prompts that users invoke as templates. For background on how agents perceive and act, see our guide to intelligent agents in AI.
Focus: Designing tool boundaries and permissions.
Tech Stack: An official MCP SDK (Python or TypeScript), SQLite, an MCP-capable client or your own agent loop.
Deliverables:
- An MCP server with at least three tools, one resource and one prompt.
- A permission model where write tools need explicit approval.
- A log of 20 agent sessions with tool calls and results.
Skills practised: Protocol design, JSON Schema, least-privilege tooling. If you'd like guided labs on tool use and MCP patterns, the Agentic AI Engineering with Claude training works through them.
What it proves to a recruiter: You can connect models to real systems safely, which sits at the core of agent engineering.
20. LLM Evaluation Harness
Build one reusable harness that runs any of your earlier projects against a test set and scores the results. Ragas provides RAG metrics including faithfulness, context precision, context recall and response relevancy, and some of them use an LLM to do the scoring.
Focus: Turning "it seems better" into numbers that block bad releases.
Tech Stack: pytest, Ragas, a YAML test set, GitHub Actions.
Deliverables:
- A test set of 50 or more cases with expected answers and source passages.
- Scored runs comparing two models or two prompt versions.
- A CI job that fails if faithfulness drops below a threshold you set and justify.
Skills practised: Evaluation design, LLM-as-judge limits, regression testing.
What it proves to a recruiter: You think like a production engineer, which is rare in junior portfolios.
21. Guardrails and Safety Layer
Wrap a chatbot with a layer on both sides: one that spots prompt injection attempts, redacts personal data such as phone and Aadhaar-style numbers, and blocks unsafe output. The OWASP Top 10 for LLM Applications 2025 makes a good checklist, covering risks such as prompt injection, sensitive information disclosure, excessive agency and unbounded consumption.
Focus: Testing attacks against your own system.
Tech Stack: Regex and a small classifier for PII, an injection detection prompt or model, FastAPI middleware.
Deliverables:
- An attack test suite of 40 or more prompts mapped to OWASP categories.
- Block rate and false positive rate on normal user questions.
- A design note on what the layer does not catch.
Skills practised: LLM security, red teaming, trade-offs between safety and usefulness.
What it proves to a recruiter: You take security seriously before someone asks.
22. Voice Agent for Appointment Booking
This assistant listens to a request, checks a mock calendar through a tool, books a slot and confirms out loud. Latency is the constraint that will bite you, so time every stage.
Focus: End-to-end latency across speech to text, model and text to speech.
Tech Stack: Whisper or a streaming speech to text service, an LLM API with tool use, a text to speech engine, WebSockets.
Deliverables:
- A recorded demo of three complete bookings.
- A latency breakdown per stage for 20 calls.
- Handling for interruptions and unclear requests.
Skills practised: Streaming, tool calling under time pressure, conversational design.
What it proves to a recruiter: You can build real-time AI that goes beyond simple request and response.
23. Pull Request Code Review Agent
Every pull request in your own repository triggers the agent. It reads the diff, runs linters and tests as tools and posts comments that point to specific files and lines.
Focus: Keeping comments useful and avoiding noise.
Tech Stack: GitHub Actions, the GitHub REST API, an LLM API with tool use, ruff and pytest as tools.
Deliverables:
- A working workflow file and agent code.
- Ten seeded pull requests with known bugs and a detection rate.
- A comment acceptance log: how many suggestions you actually applied.
Skills practised: Agents inside CI, diff context, precision of feedback.
What it proves to a recruiter: You can fit an agent into an existing engineering workflow.
24. Self-Correcting Agentic RAG
Take project 9 further. The system grades the chunks it retrieved, rewrites the query when they're irrelevant and declines to answer after two failed tries.
Focus: Reflection loops with a hard stop.
Tech Stack: LangGraph, your pgvector store, a grader prompt, the evaluation harness from project 20.
Deliverables:
- A graph with retrieve, grade, rewrite and answer nodes.
- Wrong-answer and "I don't know" rates before and after.
- Added latency and token cost per query.
Skills practised: Agentic control flow, abstention, cost and quality trade-offs.
What it proves to a recruiter: You know when an agent loop earns its extra cost.
25. Cost-Aware LLM Router with Caching
A gateway sends easy requests to a cheaper model and hard ones to a stronger model, caches repeated requests and records tokens, latency and cost for each one.
Focus: Running LLM features like a service with a budget.
Tech Stack: FastAPI, Redis or SQLite for caching, two model providers or a hosted and a local model, a simple dashboard.
Deliverables:
- Routing rules and a classifier prompt or heuristic.
- A dashboard of requests, cache hit rate and cost per day.
- An evaluation showing quality did not drop on routed requests.
Skills practised: LLM operations, caching, observability.
What it proves to a recruiter: You can keep an AI feature affordable after launch.
How to Structure Generative AI Projects on GitHub
Browse generative AI projects on GitHub and you'll find a lot of single-notebook repos, so a clean structure is an easy way to stand out:
pdf-chat-assistant/
README.md
.env.example # variable names only, never real keys
pyproject.toml
src/
ingest.py # load, chunk, embed
retrieve.py
generate.py
app.py # Gradio or FastAPI entry point
prompts/
answer_v1.txt
answer_v2.txt
eval/
test_set.csv # question, expected answer, source page
run_eval.py
results/
data/
README.md # where the data came from and its licence
Dockerfile
README checklist:
- One sentence on the problem and who it is for.
- A screenshot or 30 second demo GIF.
- Architecture diagram: data in, retrieval, model, output.
- Setup in three commands or fewer.
- Evaluation results table, including failures.
- Cost per 100 queries from your own token logs.
- Data sources and licences.
- Known limitations and next steps.
When you publish generative AI projects with source code, keep API keys out of the repository: commit .env.example, and add .env to .gitignore before your first commit.
Evaluating and Deploying Your Generative AI Project
Evaluating a project means checking its answers against expected results on a fixed test set whenever something changes. Deploying it means anyone can use it from a link, with limits in place so it doesn't run up your bill or leak data.
| Check | What to measure | Simple method | Advanced method |
| Answer quality | Correct, partly correct, wrong | Hand-score 20 to 50 cases | LLM judge plus spot checks |
| Grounding | Does the answer stick to sources? | Compare answer with retrieved text | Ragas faithfulness |
| Retrieval | Did the right chunk come back? | Hit rate at 5 | Context precision and recall |
| Latency | Seconds to first token and to full answer | Timer around each call | Per-stage tracing |
| Cost | Tokens in and out per request | Log usage fields from the API | Daily cost dashboard |
| Safety | Injection and PII handling | 20 attack prompts | Full OWASP-mapped suite |
A Gradio app on a free hosting tier is fine for a beginner demo. Once real users are involved, move to a FastAPI backend in a Docker container, add per-user rate limits, cap max_tokens on every call and set a monthly spending limit in your provider's console. Log prompts and responses too, but strip personal data before you store them.
How to Present Generative AI Projects on Your Resume and in Interviews
Recruiters skim, so lead each project line with the outcome and your measured number, then the stack. Compare these two:
- Weak: "Built a RAG chatbot using LangChain and OpenAI."
- Strong: "Built a cited Q&A assistant over 120 policy PDFs; hybrid retrieval raised hit rate at 5 from N% to M% on a 40 question test set (pgvector, FastAPI)."
Replace N and M with results you actually measured, never estimates. In interviews, expect three questions about any gen AI projects on your resume: how you know it works, what it costs per request, and what happens when retrieval returns nothing useful. Have a two minute walkthrough ready that covers the problem, the architecture, one failure you found and how you fixed it.
Common Mistakes in Generative AI Projects
- No test set. Without one, every prompt change is a guess. Write 20 questions before you write the app.
- Blaming the model for retrieval failures. Check what was retrieved before rewriting the prompt.
- Putting API keys in notebooks. Keys pushed to a public repo can be found and misused quickly. Rotate any key you've exposed.
- Giving agents write access too early. Start read-only, put write tools behind approval and log everything. OWASP lists excessive agency as a top risk for good reason.
- Ignoring cost. A loop that calls the model ten times per question can cost ten times more than you planned. Log tokens from the start.
- Copying a tutorial repo unchanged. Many LLM projects on GitHub look identical. Change the data, the problem or the evaluation so the work is clearly yours.
Conclusion
Pick one beginner project, finish it with a test set and a README, and push it this week. Then climb one level at a time. Three finished builds, say a PDF chat assistant, a measured multi-document knowledge base and an MCP tool-using agent, tell a clearer story about your generative AI projects than ten half-built demos ever will.
When you're ready to design these systems end to end, from retrieval and agents through to evaluation and deployment, the Generative AI Architect advanced program by Simpliaxis follows that same progression.























