All Tags

Tag

rag

7 items tagged with "rag"

Course

AI Hero: 7-Day AI Agents Crash-Course

Build a complete AI agent from data ingestion to deployment in 7 focused days. You will use a real GitHub repository as the source material, then build an assistant that can read the project, search through its docs and code, answer questions with citations, and ship as a small web app. This is not a toy chatbot course. By the end, you will have a portfolio-ready AI project with code, evaluation, a demo, and a clear README. What your agent will be able to do Your agent will understand a real codebase, not a tiny example file. You can use your own repository, or choose an open-source project if you do not have one ready. During the course, you will build an assistant that can: Read documentation, code, issues, and comments from a GitHub repository Split large files into useful chunks for retrieval Search with lexical search, semantic search, and hybrid search Answer project-specific questions using RAG Use tools and function calling for more useful responses Keep answers grounded with citations and logs Evaluate the quality of search and agent responses Run behind a simple Streamlit interface Be shared as a portfolio project Example question: How does authentication work in this project? By the end of the course, your agent should be able to answer questions like this by looking at the project files it indexed and pointing you to the relevant parts of the codebase. What you need Plan for 1-2 hours per day. The lessons are practical and build on each other, so it is better to do a little every day than to read everything at once. You will need: Basic Python skills: functions, classes, virtual environments, and pip Comfort with the command line A GitHub repository to analyze, or an open-source repository you want to study An OpenAI API key for the hands-on exercises A laptop where you can run Python locally The OpenAI API cost for the course is usually small. For most projects, expect a few euros or dollars, depending on the size of the repository and how much you experiment. What you will build in 7 days Each day adds one piece to the same project. Day 1: ingest and index your data from GitHub Day 2: chunk and prepare the data for search Day 3: add lexical, semantic, and hybrid search Day 4: turn the search system into an agent with tools Day 5: add logging, citations, and evaluations Day 6: publish the agent with a Streamlit UI Day 7: write the README, record a demo, and share the result Certificate: submit the finished project and get the completion certificate The syllabus below contains the detailed lessons and homework for each day. What you will have at the end When you finish, you will have a working AI assistant that understands a GitHub project and can answer questions about it. You will also have the supporting pieces that make it useful outside a notebook: Source code you can keep extending A documented ingestion and search pipeline An agent with tools and prompts Evaluation data and checks A deployed web interface A README with screenshots or a demo video A certificate of completion after project submission The goal is to leave with something you can show: in a portfolio, in a job conversation, or as a base for your next AI engineering project. Frequently asked questions What makes this different from other AI tutorials? You build with a real GitHub project, not a generic demo. The course connects the full path from data ingestion to search, agent behavior, evaluation, deployment, and portfolio packaging. What if I do not have a GitHub project? Use an open-source project that interests you. The important part is learning with real code and real documentation. Do I need an OpenAI API key? Yes. The hands-on exercises use the OpenAI API. The course shows how to keep usage small while still building the full workflow. Can I do it on a phone or tablet? You can read the lessons anywhere, but you need a computer for the coding exercises. Why is it free? This course gives you a practical foundation and shows the teaching style we use at AI Shipping Labs. If you want more structure, feedback, and accountability after that, join the community and continue with sprints, workshops, and project reviews.

May 6, 2026
Event

From RAG to Agents: Implementing Agentic Search

We build a classic RAG system over real documentation, then evolve it into an agentic search workflow where the LLM decides what to search for and whether to open a full document. Along the way we implement two tools - search (returns highlighted snippets) and get_file (returns the full document) - and wire them into an agent using both toyaikit and PydanticAI. The knowledge base is the Evidently AI documentation: a real, evolving set of Markdown files that makes RAG genuinely useful, since LLMs can't keep up with library docs on their own. This workshop was originally delivered at DataMakersFest 2026 in Porto, Portugal. There is no recording available. The system you will build The architecture has one LLM agent with two tools and two data sources: flowchart LR USER["User"] AGENT["Agent (LLM)"] SEARCH["search tool<br/>highlighted snippets"] GETFILE["get_file tool<br/>full document"] MINSEARCH["minsearch index<br/>Evidently docs"] FILEINDEX["file_index<br/>filename -> content"] USER -->|question| AGENT AGENT -->|decides which tool| SEARCH AGENT -->|decides which tool| GETFILE SEARCH -->|query| MINSEARCH MINSEARCH -->|snippets| SEARCH GETFILE -->|filename| FILEINDEX FILEINDEX -->|full text| GETFILE SEARCH -->|results| AGENT GETFILE -->|results| AGENT AGENT -->|answer| USER The agent has two tools and decides which one to use. search returns short highlighted snippets so the agent can decide which documents are worth reading. get_file returns the full document so the agent can read it end-to-end. The pattern mirrors how humans read documentation: search, scan snippets, open the promising one. Links Resources not included in the workshop materials list: toyaikit - teaching framework for agents PydanticAI - production agent framework

May 4, 2026
Event

Build a Course FAQ Agent with MCP, PydanticAI and OpenAI

We build a course FAQ assistant from the bottom up. First we expose a plain Python search(query) function to the OpenAI Responses API. Then we turn the same idea into a reusable agent loop and compare toyaikit, OpenAI Agents SDK, and PydanticAI. Finally we move the FAQ tools behind an MCP server. From there a notebook, PydanticAI, Cursor, and VS Code can all reach them. Links The main resources: AI Bootcamp: From RAG to Agents Prerequisite workshop: Building a Coding Agent Data Engineering Zoomcamp FAQ source document Parsed FAQ JSON FAQ parsing notebook The system you will build The final setup looks like this: flowchart LR NOTEBOOK["Jupyter notebook"] OPENAI["OpenAI Responses API"] FRAMEWORKS["Agents SDK<br/>PydanticAI"] MCPCLIENT["MCP clients<br/>toyaikit, PydanticAI, Cursor"] MCPSERVER["FastMCP server<br/>SSE or stdio"] TOOLS["FAQ tools<br/>search, add_entry"] INDEX["minsearch index<br/>FAQ JSON"] NOTEBOOK -->|function calling| OPENAI NOTEBOOK --> FRAMEWORKS FRAMEWORKS -->|tool calls| TOOLS NOTEBOOK -->|MCP client| MCPCLIENT MCPCLIENT -->|MCP protocol| MCPSERVER MCPSERVER --> TOOLS TOOLS --> INDEX The FAQ data comes from the Data Engineering Zoomcamp FAQ. The first half of the workshop keeps the tools inside the notebook so you can see the agent loop directly. The second half moves the same tools into mcp_faq/, which makes them reusable by any MCP client.

Sep 1, 2025
Event

Build Your Own Search Engine

We build a search engine from scratch over DataTalks.Club Zoomcamp FAQ documents. We start with TF-IDF text search, then add cosine similarity and field boosting. From there we move through SVD/LSA and BERT embeddings to vector search. The TextSearch class built during the workshop became the basis for the minsearch library used in later RAG and agent workshops. What we cover: TF-IDF text search with sklearn, cosine similarity, field boosting and keyword filtering A reusable TextSearch class that becomes minsearch Vector search using SVD and NMF embeddings BERT embeddings for semantic search that respects word order Originally delivered at a DataTalks.Club live session in 2024, updated in 2026 with refreshed examples and tooling. Links Resources not included in the workshop materials list: DataTalks.Club YouTube channel The search engine you will build We take two approaches to search over the same FAQ data: flowchart LR DOCS["FAQ documents<br/>DE/ML/MLOps Zoomcamp"] TEXT["Text search<br/>TF-IDF + cosine similarity<br/>field boosting + filtering"] VEC["Vector search<br/>SVD / NMF / BERT embeddings<br/>cosine similarity"] CLASS["TextSearch class<br/>(became minsearch)"] DOCS --> TEXT DOCS --> VEC TEXT --> CLASS VEC --> CLASS Text search uses TF-IDF vectorization, weights the question field three times as much as the others, and filters by keyword. Vector search replaces sparse representations with dense embeddings (SVD, NMF, then BERT) to handle synonyms and word order. Both paths share the TextSearch class we build along the way, which combines TF-IDF across multiple fields with boost weights and keyword filters.

May 21, 2024