# Overview This tutorial shows how to build a documentation-focused chatbot that finds relevant instructions inside preindexed docs and answers follow-up questions with context. The example uses UpCloud docs, but the same pipeline applies to other documentation collections.
# How it works The app uses retrieval-augmented generation (RAG). Apify crawls and collects documentation pages and returns text plus source URLs. The text is split into passages. OpenAI turns passages into embeddings. Pinecone stores those embeddings and supports semantic search. LangChain wires these pieces together: it lets the LLM query Pinecone, call Apify for live lookups when needed, and use LangGraph to keep conversation history for follow-ups. Streamlit hosts the chat interface.
# What you need before you start
- Python 3.12 and a code editor (PyCharm suggested).
- Accounts and API credentials for OpenAI, Pinecone, and Apify. OpenAI API billing must be enabled. An Apify account needs credits if you run their crawlers.
- GitHub and Streamlit Community Cloud accounts if you want to publish the app.
# Project layout and dependencies Create a project folder (example: upcloud-docs-bot) and these files: ingest.py (collects and indexes docs), app.py (runs the chatbot UI), requirements.txt,.env (local API keys),.gitignore. The tutorial uses these pip packages (versions shown in the source): langchain and related langchain-* packages, apify-client, streamlit, python-dotenv. Install dependencies inside the project virtualenv.
# Core components and steps
- 1Split text: Divide pages into passages suitable for embedding and retrieval.
- 2Create embeddings: Use the OpenAI API to convert passages to embeddings.
- 3Index embeddings: Create a Pinecone index and store passage embeddings together with source URLs.
- 1Generate answers: Pass retrieved passages and conversation history to the LLM via LangChain to produce an answer that cites sources.
- 1Conversation state: LangGraph (installed as part of the LangChain packages) manages earlier messages so follow-up questions stay coherent.
- 2UI: Streamlit hosts the chat interface that you can run locally or publish on Community Cloud.
# Practical notes
- Store API keys in.env and add.env to.gitignore. A ChatGPT subscription does not replace OpenAI API usage. Pinecone gives a default API key on sign-in. Apify supplies an API token in integrations settings.
- The tutorial includes a downloadable GitHub repo with complete code you can clone and run with your own keys.
- Streamlit Community Cloud offers free hosting, but OpenAI, Pinecone, and Apify each have separate usage limits and charges to monitor.
# Where this is useful A documentation chatbot reduces time spent hunting for instructions, helps users get guidance that fits their setup, and supports follow-ups that depend on conversation context. It is suitable when your docs are stable enough to index but occasional live lookups are still needed.
# Quick deployment checklist
- Python 3.12 virtualenv created inside project folder.
- requirements.txt installed in that virtualenv.
- .env populated with OpenAI, Pinecone, and Apify credentials.
- Pinecone index created and configured for the chosen embedding dimension.
- ingest.py run to populate the index and namespace file produced.
- app.py run locally or deployed to Streamlit Community Cloud.
# Next steps Use the included GitHub repo to test with your credentials, adjust text-split sizes for your docs, tune Pinecone search parameters, and add or swap OpenAI model settings to balance cost and response quality.