Skip to content
Woyce Technologies
AboutTeamCareersContactStart a project →

How to Build an AI Chatbot with LangChain and OpenAI

A practical guide to building a production-ready AI chatbot using LangChain, OpenAI GPT-4, and Next.js — with RAG for grounding answers in your own data.

How to Build an AI Chatbot with LangChain and OpenAI — Woyce Technologies

Building an AI chatbot that actually works for your business takes more than calling the OpenAI API. You need conversation memory, grounding in your own data, and a sensible way to handle edge cases. LangChain gives you the scaffolding for all of that — though you'll still write the parts that matter most yourself.

In this guide, we'll build a production-ready chatbot using LangChain, OpenAI GPT-4, and Next.js with Retrieval-Augmented Generation (RAG) so your bot answers from your business's actual documents — not just whatever the model picked up in training.

The distinction matters more than most people realise. A raw GPT-4 call will answer customer questions confidently and incorrectly if the answer isn't in its training data. A 12-person law firm that deployed a plain GPT-4 chatbot for client intake found it citing statutes that had been amended two years prior — correct-sounding, completely wrong. RAG solves this by pulling from documents you control and update. The model's job becomes synthesis, not recall.

Below you'll find the dependencies to install, the RAG chain, the API route, document ingestion, a minimal chat UI, a comparison with off-the-shelf chatbot tools, what production deployments actually run into, and the mistakes that cost teams the most time.

What You'll Build: Your AI Chatbot

  • A Next.js API route that handles chat messages
  • A LangChain conversation chain with memory
  • A RAG pipeline using Pinecone as a vector store
  • A simple React chat UI

The architecture is deliberately minimal. A real deployment will need auth, rate limiting, logging, and error boundaries — but this gives you a working foundation you can extend without unpicking someone else's decisions.

Prerequisites

You don't need prior LangChain experience, but you should be comfortable with async TypeScript and Next.js API routes. If you've never used a vector database, skim the Vector databases explained post first — it'll make the Pinecone step much clearer.

Step 1: Install Dependencies

npm install langchain @langchain/openai @langchain/pinecone @pinecone-database/pinecone

Pin your versions. LangChain's API surface changes frequently across minor versions. At the time of writing, langchain@0.3, @langchain/openai@0.3, and @pinecone-database/pinecone@3.x are stable together. Mixing patch versions from different release cycles is the single fastest way to lose an afternoon to TypeScript errors that should not exist.

Step 2: Set Up the RAG Pipeline

Create src/lib/chatbot.ts to initialise your LangChain chain with Pinecone retrieval and conversation memory.

The chain uses ConversationalRetrievalQAChain, which retrieves the top 4 relevant document chunks from Pinecone, injects them into the prompt, and passes conversation history through BufferMemory.

Here is what is happening under the hood: when a user sends a message, LangChain first embeds that message using the same OpenAI embeddings model you used during ingestion. It then runs a similarity search against your Pinecone index and returns the four chunks with the highest cosine similarity. Those chunks get appended to the system prompt as context before the message hits GPT-4. The model sees both the retrieved context and the conversation history, which is why it can answer follow-up questions without losing the thread.

A few configuration choices worth making deliberately:

k: 4 — the number of retrieved chunks. Four is a reasonable default. Too few and you miss relevant context; too many and you burn tokens on noise, which degrades answer quality and raises cost. A 50-page product manual benefits from k=6. A tightly scoped FAQ chatbot often does better at k=3.

The embedding model. text-embedding-3-small is significantly cheaper than text-embedding-ada-002 and performs comparably on most retrieval tasks. Unless you have a specific reason to use ada-002, start with small.

The system prompt. This is where most projects go wrong. Do not skip it. Tell the model its role, what it should and should not answer, and how to handle questions that fall outside the retrieved context. A well-written system prompt cuts hallucinations more than any retrieval tuning.

Step 3: Create the API Route

Create src/app/api/chat/route.ts as a Next.js API route that accepts POST requests with a message field and returns the AI response.

Keep the route thin. Its job is to validate input, call the chain, and return a response — not to contain business logic. If you find yourself writing more than 60 lines here, something belongs in src/lib/chatbot.ts instead.

One pattern worth adding from the start: return the source documents alongside the answer. LangChain's ConversationalRetrievalQAChain gives you sourceDocuments on the response object. Surfacing these in your UI as citations lets users verify answers rather than simply trusting them — which matters particularly for legal, medical, or financial contexts where a wrong answer has real consequences.

Error handling is not optional. The OpenAI API returns HTTP 429 when you hit rate limits, and Pinecone times out under load. Handle both explicitly. A generic 500 response with no context makes debugging much harder at 2am when something breaks in production.

Step 4: Ingest Your Documents

Before users can chat, embed your documents into Pinecone: use RecursiveCharacterTextSplitter to chunk the text, OpenAIEmbeddings to create vectors, and PineconeStore.fromDocuments to store them. Tune the chunk size to your content — there's no universal right answer, and we've ended up adjusting it on almost every project.

What that adjustment looks like in practice: a property management company we worked with had lease agreements that ran 40–60 pages each. Their first pass used a 1000-token chunk size, which was splitting mid-clause. Retrieval was returning partial sentences, and the model was filling gaps with plausible but incorrect lease terms. Dropping to 512 tokens with a 50-token overlap fixed the clause fragmentation completely. A software documentation project ran in the opposite direction — short API reference entries were getting split across chunks, losing the pairing between a function signature and its explanation. Larger chunks (1500 tokens) and higher overlap (200 tokens) solved it.

The ingestion script should be idempotent. If you run it twice, you do not want duplicate vectors. Pinecone supports namespacing and upsert by ID — use both. Assign a stable ID to each document chunk (hash the source path plus chunk index, for example) so re-running the ingestion script updates existing vectors rather than adding duplicates.

Store metadata with every vector: source filename, page number, last modified date. You will want to filter by these during retrieval, and you cannot add metadata retroactively without re-ingesting.

ApproachOff-the-shelf chatbot (e.g. Intercom AI)Custom LangChain + RAG build
Setup timeHours2–6 weeks depending on data complexity
Data groundingLimited to public knowledge or simple FAQ uploadFull control — any document format, any update cadence
Answer accuracy on proprietary contentLow to moderateHigh when tuned correctly
Conversation memorySession-only, vendor-managedConfigurable — Redis, DB, or in-memory
Monthly cost at scale$300–$2,000+ (per-seat or usage pricing)OpenAI API + Pinecone (~$50–$400 depending on volume)
CustomisationTheme and copy onlyFull control over prompts, retrieval logic, UI
Maintenance burdenVendor-managedInternal or outsourced

Step 5: Build the Chat UI

The front end can stay small: a client component that keeps an array of messages in state, posts the latest user message plus recent history to /api/chat, and appends the reply. Resist the urge to build a full design system for it on day one. What matters is that the UI makes the bot's behaviour legible.

Three details are worth getting right early:

  • Stream the response. Render tokens as they arrive instead of waiting for the full answer. The total time doesn't change, but users stop wondering whether the bot has frozen.
  • Show sources. If the API route returns sourceDocuments, render them as small citations under each answer. It builds trust and gives your team a quick way to spot bad retrievals.
  • Make failure visible. Show a clear message when the API returns an error or a rate-limit response, and offer a route to a human or a contact form instead of a spinner that never stops.

Accessibility is cheap to add now and expensive later: give the message list an aria-live="polite" region so screen readers announce new replies, and keep the input usable by keyboard alone.

Benefits of a LangChain and OpenAI Chatbot

Off-the-shelf tools get you a chat widget quickly. Building on LangChain with RAG costs more effort up front and returns a different set of advantages.

Answers Grounded in Your Own Documents

The core benefit is that answers come from content you control. Instead of relying on whatever the model absorbed in training, every response is built from chunks retrieved from your manuals, policies, or knowledge base. When a policy changes, you re-ingest the document and the bot's answers change with it. That is what makes the chatbot usable for questions where a confident but outdated answer would do real harm, such as returns rules, contract terms, or product specifications.

Control Over Every Layer

LangChain exposes the pieces separately: the splitter, the embedding model, the vector store, the retriever, the memory, and the prompt. You can tune k, add metadata filters, swap Pinecone for pgvector, or change how history is summarised without rebuilding the whole system. Hosted chatbot products usually hide those decisions, which is fine until retrieval starts returning the wrong passages and there is nothing you can adjust.

Verifiable Answers With Citations

Because the chain returns sourceDocuments, the UI can show exactly which passages informed each reply. Users can check the claim themselves, and your team can spot a bad retrieval at a glance. For legal, financial, or technical content, that traceability is often the difference between a tool people trust and one they quietly stop using.

Costs That Track Usage

You pay for model tokens and vector storage rather than per-seat licences. With sensible chunk sizes and a modest k, the running cost of a focused chatbot stays proportional to how much it is actually used. Optimising retrieval reduces both token spend and noise in the prompt, so cost work and quality work tend to point in the same direction.

Portability Across Models and Stores

LangChain's abstractions make it practical to move between OpenAI models, or to another provider, and between vector databases, with contained code changes. That reduces lock-in on two of the most fast-moving parts of the stack. Your document pipeline, metadata, and evaluation questions remain your own assets regardless of which model sits underneath.

LangChain Chatbot Use Cases

The same architecture fits a range of jobs. What changes between them is the document set, the metadata you filter on, and how strict the system prompt needs to be.

Customer Support Over Help Content

Support teams field the same questions about orders, accounts, and policies every day, and the answers already exist in help articles. A RAG chatbot ingests that content and answers directly, with links to the source article. The e-commerce example above shows the main work: deciding which documents to include and excluding outdated policies. Done well, routine questions get accurate answers instantly and the team handles the exceptions.

Internal Knowledge Assistants

Staff waste time searching wikis, shared drives, and old threads for procedures or past decisions. An internal chatbot over curated internal documents lets them ask in plain language and get an answer with the source attached. The recruitment firm's candidate search shows the key lever: metadata filters, so a question about a specific domain retrieves only the relevant records rather than anything sharing a generic keyword. Access control matters here too, so restrict retrieval to documents each user is allowed to see.

Product and API Documentation

Developer and product documentation is long, versioned, and often split into many short reference entries. A documentation chatbot retrieves the relevant sections and explains them in context. As the earlier chunking example showed, short reference entries need larger chunks with more overlap so function signatures stay with their explanations. Users find answers without paging through the docs, and the docs team learns which questions the content isn't answering. Unanswered questions logged by the bot become a practical to-do list for the documentation team.

Long Contracts and Policy Documents

Leases, contracts, and policy manuals are dense, and people usually need one clause rather than the whole document. A chatbot over these files, with smaller chunks so clauses aren't split, can locate and summarise the relevant terms and cite the page. The system prompt should tell it to quote rather than paraphrase where precision matters, and to defer to a person for interpretation. The result is faster lookups with a clear path back to the original text.

What to Expect in Practice

The first working version typically takes one to two days to get running locally. Getting it production-ready is a different scope.

A mid-size e-commerce brand with 8,000 SKUs and 5 years of customer support transcripts took three weeks to go from proof-of-concept to production. The technical build was the smaller part. Most of the time went into deciding which documents to ingest (and which to exclude — outdated return policies created more confusion than no policy at all), writing the system prompt through iteration, and setting up logging so the team could monitor what the bot was getting wrong.

A recruitment firm used a similar stack to build an internal chatbot over their candidate database and job specs. Their main challenge was retrieval relevance — a query for "senior engineer with payments experience" was matching on "engineer" and "experience" but missing the "payments" context. Adding metadata filters on industry tags at the Pinecone query stage fixed the accuracy meaningfully.

Response latency is worth thinking about early. A RAG pipeline with GPT-4 typically takes 2–5 seconds end-to-end: embedding the query, running the vector search, and generating the response. For a support chatbot this is acceptable. For an internal tool used dozens of times per hour by staff, it adds up. Streaming the response (OpenAI supports SSE) makes the perceived latency much shorter even when the total time is the same.

Common LangChain Chatbot Mistakes

Using BufferMemory in Production Without Persistence

This is the most common oversight. BufferMemory holds conversation history in the Node.js process. When the process restarts — every deployment, every server event — the history is gone. Users mid-conversation lose all context, and on serverless hosting it can disappear between requests. Use Redis or a database-backed memory store from the start.

Ingesting Everything Without Curation

More data is not always better. Ingesting five years of internal Slack messages alongside product documentation creates retrieval noise. A chatbot asked about refund policy should not be retrieving a three-year-old thread about office snacks. Be deliberate about what goes in, and remove superseded versions of documents rather than leaving old and new side by side.

Forgetting to Update the Index

Documents change. If your Pinecone index is a snapshot from six months ago, your chatbot is answering from stale data, and nobody notices until a customer quotes the old answer back. Build the ingestion pipeline as a scheduled job, not a one-time script, and make it idempotent so re-runs update vectors instead of duplicating them.

Not Testing Adversarial Queries

Ask your chatbot questions it should not answer. Questions outside its domain, leading questions, attempts to get it to make commitments it should not make, and prompts that try to override the system instructions. Find these before your users do, and keep the failing examples as regression tests. Review refusals as well: a bot that declines too often is failing in a quieter way.

Skipping Structured Logging

You will not know what's failing without logs. At minimum, log the user message, the retrieved chunk IDs, and the final response. This data is also how you improve the system over time — it's the only way to know whether your retrieval is working or whether the model is improvising around poor context.

LangChain Chatbot Best Practices

  • Pin dependency versions together. Lock langchain, @langchain/openai, and the Pinecone client to versions known to work as a set, and upgrade them deliberately in one change with tests, rather than letting minor releases drift independently.
  • Write the system prompt before tuning retrieval. Define the bot's role, its scope, and exactly what to say when the retrieved context doesn't answer the question. This does more to limit made-up answers than adjusting k or chunk size.
  • Tune chunking to the documents, not a default. Inspect what retrieval actually returns for real questions. Shrink chunks when clauses are being split mid-sentence; enlarge them, with more overlap, when short reference entries lose their context.
  • Store metadata on every vector. Source file, page, section, and last-modified date make filtering, citations, and targeted re-ingestion possible. Adding them later means re-embedding everything.
  • Keep a test set of real questions. Collect 30 to 50 questions users actually ask, with the expected answer and source, and run them after every prompt, chunking, or model change to catch regressions.
  • Persist memory and bound its length. Store history in Redis or a database, and trim or summarise older turns so long conversations don't inflate token costs or crowd out retrieved context. Decide upfront how long history is kept and who can read it.
  • Return and display sources. Show citations with each answer so users can verify, and so your team can see at a glance when retrieval went wrong. Link each citation to the original document and page where possible.
  • Plan for failure modes. Handle rate limits and timeouts explicitly, stream responses, and provide a fallback such as keyword search or a human contact when the model API is unavailable. Test each fallback on purpose before launch rather than waiting for the first outage.

Key Takeaways

  • RAG grounds your chatbot in real business data, which is the single biggest defence against hallucinated answers
  • BufferMemory gives the chatbot conversation history within a session — fine for dev, not enough for production
  • For production, store conversation history in Redis or a database rather than in-memory, or you'll lose context every time the process restarts
  • Woyce Technologies builds custom LangChain chatbots — contact us if you'd like to talk through what's right for your use case

Frequently Asked Questions

How much does it cost to run a LangChain chatbot in production?

At moderate volume (roughly 10,000 queries per month), expect $30–$120/month in OpenAI API costs depending on message length and GPT-4 vs GPT-4o. Pinecone's free tier handles up to 1 million vectors, which is sufficient for most small business document sets. The biggest cost variable is how many tokens your retrieved chunks add to each prompt — optimising chunk size and k directly reduces spend.

Do I need a vector database, or can I use a simpler approach?

For fewer than 50 short documents, you can get away with loading them all into the prompt at query time — it's simpler and has no infrastructure to manage. Beyond that, retrieval becomes necessary both for accuracy (the context window has limits) and cost (sending every document with every message gets expensive fast). Pinecone is the easiest managed option; pgvector is worth considering if you already run PostgreSQL and want to keep your stack simple.

How long does it take to build a production-ready AI chatbot with LangChain?

A focused team with prior LangChain experience typically needs 3–6 weeks for a production deployment: 1 week for the core build and ingestion pipeline, 1–2 weeks for prompt tuning and retrieval accuracy work, and 1–2 weeks for productionising (auth, logging, error handling, deployment). If your data requires significant cleaning or your use case has compliance requirements, add time.

Can this chatbot handle multiple languages?

GPT-4 handles multilingual input and output well out of the box. The retrieval step is the limiting factor — if your Pinecone index contains only English documents, a query in French will still find relevant chunks (OpenAI embeddings are multilingual), but the retrieved context will be in English, which affects answer quality. For genuinely multilingual deployments, store documents in each target language separately and route queries to the matching namespace.

How do I prevent the chatbot from making things up?

Three things work together: a tight system prompt that instructs the model to say "I don't know" when the retrieved context doesn't cover the question, returning source documents alongside answers so users can verify, and logging responses so you can identify and fix patterns of hallucination over time. RAG reduces hallucination significantly compared to a bare API call, but it does not eliminate it — the model can still interpolate beyond what the retrieved chunks say.

Is LangChain the right choice, or should I use the OpenAI Assistants API instead?

The OpenAI Assistants API is simpler to get started with and handles file retrieval natively. LangChain gives you more control: over the retrieval logic, the vector store, the memory implementation, and how prompts are constructed. If you need to integrate multiple data sources, run custom retrieval logic, or avoid vendor lock-in on your document storage, LangChain is the better foundation. If you want something working in a day and your use case is straightforward, the Assistants API is worth evaluating first. See our detailed comparison for a side-by-side breakdown.

What happens if the OpenAI API goes down?

Without a fallback strategy, your chatbot goes down with it. For production deployments, implement a graceful degradation path: a clear error message to users, a fallback to a simpler keyword-search over your documents, or routing to a human. OpenAI's API has a published SLA and status page — monitoring it and alerting on degraded performance is worth setting up from day one rather than finding out from users.

Conclusion

The hard part of building an AI chatbot with LangChain and OpenAI isn't getting a reply from the model. It's making sure the reply comes from your documents, survives restarts and deployments, and fails safely when something breaks. RAG, persistent memory, and good logging are what separate a weekend demo from a tool your customers or staff can rely on.

The most useful lessons from this guide are about data and operations rather than code. Curate what you ingest, tune chunk size to your actual documents, store metadata with every vector, and run ingestion as a scheduled, idempotent job. Write the system prompt deliberately, return sources with every answer, and log enough to see which retrievals are failing.

Keep the limits in mind too. RAG reduces hallucination but doesn't eliminate it, LangChain's API moves quickly enough that version pinning is mandatory, and API outages need a fallback path. For simpler use cases, a hosted option like the Assistants API may be the faster route.

As a next step, get the pipeline running locally against 20 to 50 of your real documents and test it with the questions your users actually ask. When you're ready to take it to production, our LLM integration team can help with retrieval tuning, memory, and deployment.

WT

Woyce Technologies

AI & Engineering Team · Woyce

Woyce Technologies builds AI chatbots, LLM integrations, voice AI, and full-stack web applications for businesses in the US, UK, Europe & APAC. Based in Rajkot, Gujarat.

READY TO BUILD?

Let's build something
that actually works.

Tell us about your project. We'll be honest about whether we're the right fit — and if we are, we move fast.