Why LLMs Struggle With Large Codebases: The Context Problem

LLMs can help with individual code files, but they often struggle with large codebases because raw code is not structured context. Learn why dependency graphs and PViz bundles give AI tools a better map of your software.

•21 min read
LLMsdependency graphsAI code analysiscodebase contextstatic analysisLLM workflowsrepository analysis

If you have ever pasted a file into Claude, ChatGPT, or another LLM and asked it to help you make a code change, you already know the experience can work surprisingly well.

For a small, self-contained task, the model reads the file, understands the local logic, and gives you something useful. Maybe it explains a function. Maybe it finds a bug. Maybe it rewrites a class or helps you add a test. When the task fits inside one file, the interaction feels almost natural.

The trouble starts when the change does not fit inside one file.

Most real software changes do not.

Adding request logging might touch your application entry point, middleware registration, route handlers, configuration loader, and shared logging utilities. Adding rate limiting might involve authentication, API routes, cache configuration, deployment settings, and error handling. Refactoring a helper function might look local until you realize it is imported by a dozen modules that rely on slightly different behavior.

The LLM does not know any of that unless you tell it.

And telling it is the part developers are still doing manually: finding the relevant files, deciding what context matters, copying code into the prompt, explaining the directory structure, and hoping the model understands how the pieces fit together.

This is the context problem.

The issue is not simply that the model is bad at code. The issue is that the input you give it often does not describe the system. It describes a fragment of the system.

That distinction matters. LLMs can be very good at reasoning over code when they have the right context. But a large codebase is not just a collection of files. It is a network of relationships: imports, dependencies, entry points, shared utilities, test boundaries, cycles, and architectural seams. If those relationships are missing from the prompt, the model has to infer them from incomplete evidence.

Sometimes that inference is good enough.

Often it is not.

This post is the first in a series about making LLM-assisted codebase analysis more grounded. The goal is not to claim that static analysis replaces reading code, or that a dependency graph gives an LLM everything it needs. It does not. The goal is narrower and more practical: before asking an LLM to reason across a repository, give it structured context about how that repository is connected.

That is the gap PViz is built to address.

PViz generates a structured JSON bundle from static analysis of a codebase. The bundle gives an LLM a map of the repository: the modules that exist, the dependency edges between them, and a folder index that summarizes where things live. Instead of asking the model to discover structure from a raw file dump, you can give it structure up front.

That shift — from raw code context to structured code context — is the foundation for this series.

The one-file illusion

LLMs feel strongest when the task is local.

If you paste in a single Python file and ask, "What does this function do?" the model can usually answer. If you ask, "Can you clean this up?" it can often produce a reasonable refactor. If you ask, "Why is this condition never true?" it may spot the logic problem quickly.

That creates what I think of as the one-file illusion.

The illusion is that because the LLM understands one file well, it understands the codebase well enough to make system-level recommendations.

But those are different tasks.

Understanding one file means reading the code that is directly visible. Understanding a codebase means understanding how that file participates in a larger system. Who imports it? What does it import? Is it an entry point, a utility, a leaf module, or a central dependency? Is it part of a cycle? Are there tests that target it directly? Does it sit on a critical path that many other modules depend on?

Those questions are not always answerable from the file itself.

A module can look small and harmless in isolation while being structurally important. A helper can look generic while serving as a hidden dependency for half the project. A route handler can look like the right place to add behavior when the better location is middleware or shared request processing.

The model can only reason from what it sees.

If it sees one file, it gives you a one-file answer.

That may be useful, but it is not the same as a codebase-aware answer.

Real changes cross file boundaries

This is where the problem becomes practical.

Developers rarely ask LLMs only theoretical questions. We ask things like:

  • "Where should I add this feature?"
  • "What files need to change?"
  • "What could this refactor break?"
  • "Which tests should I run?"
  • "Is this module safe to modify?"
  • "Why is this behavior happening across multiple layers?"

Those questions are structural.

They are not just about syntax or local logic. They are about relationships.

Consider a FastAPI backend. If you ask an LLM, "Where should I add request logging middleware?" and you only provide one route file, the model has to answer from general FastAPI knowledge. It may say to add middleware near application startup. That is reasonable. It might even be correct.

But it is still a convention-based answer.

It does not know your actual application layout. It does not know whether your app is created in main.py, an application factory, a bootstrap module, or a deployment-specific entry point. It does not know whether middleware registration is centralized or scattered. It does not know whether request IDs already exist. It does not know whether logging configuration is loaded from environment settings, a config module, or a shared observability package.

Without repository structure, the model can only make an educated guess.

That is the difference between plausible and grounded.

A plausible answer sounds right because it matches common patterns.

A grounded answer is tied to the actual structure of your codebase.

For small tasks, plausible may be enough. For larger changes, plausible can be risky.

More context is not always better context

The obvious workaround is to paste more files.

If one file is not enough, paste five. If five are not enough, paste the directory tree. If that still does not work, use a tool that concatenates the whole repository into one giant text file and feed as much of it as possible into the model.

This is understandable. Developers are trying to solve a real problem: the model needs more context.

But more context and better context are not the same thing.

A large context window helps, but it does not remove the need to decide what information matters. Even if a model can technically accept a very large input, the prompt still has to be useful. A pile of undifferentiated source files forces the model to do several jobs at once: identify the relevant files, infer the dependency relationships, determine which modules are central or peripheral, separate local implementation detail from architectural structure, and then answer the actual question.

That is a lot to ask from one prompt.

The larger the input becomes, the easier it is for important details to get buried. A relevant import path may be thousands of lines away from the question. A critical utility may appear in the middle of a file dump with no signal that it is widely used. A module may look unimportant because its code is short, even though many other modules depend on it.

So pasting more files does not really solve the problem. It changes the failure mode.

Instead of missing context, the model now has too much undifferentiated context.

That can produce a different kind of bad answer: one that appears comprehensive, references multiple files, and still misses the structural point.

A file dump is not a map

There are tools that flatten a repository into a single text blob. These can be useful. They make it easier to collect files, preserve file paths, and move code into a prompt. For some use cases, that is enough.

But a file dump is still not a map.

A file dump tells the model what code exists. It does not directly tell the model how the code is connected.

That distinction is important.

Imagine trying to understand a city from a stack of building descriptions. You could read about every building one by one: its address, height, purpose, and layout. Eventually, you might infer something about the city. But you would still be missing the road network.

The road network changes everything.

It tells you which places are connected, which areas are central, which routes are bottlenecks, and which neighborhoods are isolated. Without that map, you may know a lot about individual buildings and still misunderstand how the city works.

Codebases have the same problem.

Files are not independent objects. They import each other. They depend on shared configuration. They form layers. They create cycles. They expose entry points. They concentrate risk around heavily reused modules.

A flat dump hides those relationships in plain sight.

The import statements are technically present, but the structure is not summarized. The model has to reconstruct it from raw text. In a large repository, that reconstruction is exactly the work you wanted help with.

What you want to hand an LLM is not just a transcript of your codebase.

You want to hand it a map of the relationships inside the codebase.

What dependency graphs add

A dependency graph represents a codebase as nodes and edges.

The nodes are modules or files.

The edges are dependency relationships.

If api.routes.users imports api.services.auth, there is an edge from the route module to the auth service module. If several parts of the system import the same config module, that config module becomes structurally visible as a shared dependency. If two modules import each other directly or indirectly, that relationship can appear as a cycle.

This graph view answers questions that raw source text does not answer cleanly: which modules does this file depend on, which modules depend on this file, what files are upstream or downstream of a proposed change, which modules are highly connected, which files look isolated, where are the circular dependencies, and which directories contain most of the dependency activity.

These are the questions developers ask before changing real systems.

They are also the questions an LLM needs answered before it can provide useful codebase-level guidance.

A dependency graph does not replace source code. It does not tell you every runtime behavior. It does not prove that a change is safe. Static analysis has limits, especially in dynamic languages and frameworks where behavior can be created at runtime.

But the graph gives the model something raw code does not: orientation.

It helps the model understand where to look before it starts giving advice.

That orientation is often the difference between a generic recommendation and a useful one.

The context problem is really a selection problem

When developers talk about context windows, the conversation often turns into a question of size: how many tokens can the model accept, how many files can I fit, can I paste the whole repository?

Those are reasonable questions, but they can obscure the deeper issue.

The real problem is selection.

For a given task, which parts of the codebase matter?

If you are changing a central configuration loader, you need to know who depends on it. If you are modifying a route handler, you may need its service layer, schemas, tests, and middleware path. If you are refactoring a shared utility, you need to understand downstream callers. If you are reviewing architecture, you may need hubs, cycles, and directory-level structure.

Different tasks require different slices of context.

A full repository dump treats every file as equally relevant. A dependency-aware view can help narrow the question.

This does not mean the model should only receive graph data. It means graph data can help decide which source files are worth reading next.

That is a more practical workflow: use structured analysis to orient the model, identify the relevant modules and dependency paths, bring in source code for the files that actually matter, then ask the model to reason from both structure and implementation.

That workflow is much closer to how a developer investigates a codebase manually.

You do not usually open every file at random. You start with an entry point, trace dependencies, inspect related modules, and narrow the problem.

A structured bundle helps the LLM follow the same pattern.

How PViz turns a codebase into LLM context

PViz generates a structured JSON bundle from static analysis of a repository. The purpose of the bundle is to give an LLM compact, machine-readable context about the shape of the codebase.

At a high level, the bundle contains three major categories of information.

Nodes. A node represents a module or file in the codebase. For each node, the bundle includes the file path, language, declared imports, exported symbols, and graph-related metadata. The point is to give the model an inventory of what exists in the repository without requiring it to read every line of source code first.

A simplified node looks like this:

{
  "node": "api/main.py",
  "imports": ["api.routes.users", "api.middleware.logging"],
  "role": "entrypoint"
}

Edges. Edges describe dependency relationships between nodes. If one module imports another internal module, that relationship becomes an edge in the graph. These edges are what allow the model to move from "this file exists" to "this file depends on that file" and "this file is depended on by these other files." Edges are central to change-impact questions — if you are changing a module, the edge data helps identify related modules that may need inspection.

Folder index. The folder index gives a hierarchical view of the project structure. This helps the model orient itself before drilling into individual nodes and edges. Directory names often carry architectural meaning: api, services, models, workers, tests, middleware, config. The folder index gives the model a compact overview of that layout before it starts traversing the graph.

Together, nodes, edges, and the folder index provide a structured map of the codebase. They do not replace reading source files. They help decide what source files should be read.

Why this matters for AI code review

AI code review is often framed as a question of whether the model can understand code.

But for codebase-level review, the harder question is whether the model understands impact.

A local code review asks whether a function is correct, whether an implementation is readable, whether there are obvious bugs, or whether something could be simplified.

A codebase-aware review asks harder questions: what depends on this module, does this change affect an entry point, is this function part of a cycle, are there downstream modules that assume the old behavior, are there tests near the affected dependency path, is this change touching a hub module used throughout the project?

Those questions require structure.

This is why dependency graphs are useful for LLM code analysis. They turn invisible relationships into explicit context.

Without that structure, the model may review a change as if it were local when it is actually systemic.

That is one reason LLMs can feel inconsistent on large repositories. The same model that gives excellent feedback on one file can miss the broader risk of changing that file.

The problem is not only intelligence. It is visibility.

Example: adding request logging middleware

Suppose you ask an LLM:

I want to add request logging middleware. What files will I need to touch?

With only a generic prompt, the model will likely answer based on common FastAPI conventions. It may suggest looking for main.py, adding middleware with app.middleware("http"), configuring logging, and possibly updating tests.

That is not a bad answer.

But it is not grounded in your repository.

With a raw file dump, the model may have more information. It may see your route files, config files, and app startup code. But unless the repository is small, it still has to infer which files matter from a large amount of undifferentiated text.

With a PViz bundle, the prompt starts differently.

Instead of asking the model to guess where the application starts, you give it the graph structure and ask it to identify likely entry points, middleware-related modules, logging utilities, and downstream route modules based on actual dependency relationships.

The answer is then framed around the repository's structure: the application entry point or app factory, the module where middleware is registered, configuration modules used by the app startup path, logging-related utilities if they exist, route modules downstream of the app registration path, and tests that appear near the affected modules.

The model still needs source code before writing the final implementation. But it now has a better starting point.

It can say, "These are the files I would inspect first, and this is why."

That is a more useful answer than, "In a typical FastAPI app, you might…"

The goal is not to make the model sound confident. The goal is to make its confidence easier to audit.

Why not just use RAG?

Retrieval-augmented generation is another common approach to the context problem. In a RAG workflow, the system retrieves chunks of relevant code or documentation and gives those chunks to the model.

That can be useful, especially when the question is text-based: "Where is this function defined?" or "Find files related to authentication."

But codebase analysis is not only text retrieval.

A file can be relevant because of its position in the graph, not because it contains the same words as the question.

A central configuration module might not mention "request logging" at all, but it may still matter because application startup depends on it. A shared utility might not match the query terms, but it may be imported by every route. A test file might not use the same language as the feature request, but it may sit next to the affected module.

Keyword relevance and structural relevance are different signals.

The best workflow uses both. Text retrieval helps find content. Graph structure helps explain relationships. PViz sits on the structural side of that equation — it is not a replacement for retrieval, but an orientation layer that makes retrieval more targeted. In Part 5, we will look at how to combine the two into a full codebase Q&A pipeline.

From plausible answers to grounded answers

A plausible answer is based on patterns the model has learned.

A grounded answer is constrained by evidence from your repository.

Plausible answers are not useless. If you ask about a common framework, the model's general knowledge can be valuable. But when the question depends on your specific codebase, general knowledge is not enough.

The model needs evidence.

A dependency graph provides one form of evidence. It tells the model what files exist and how they relate. That makes it harder for the model to invent architecture that is not present, and easier for you to evaluate whether the answer follows from the actual repository.

This matters especially for larger codebases, older codebases, and codebases that do not follow textbook structure.

Real projects accumulate exceptions. They contain legacy modules, migration paths, utilities with unclear ownership, tests that only cover part of the system, and naming conventions that made sense three years ago. A model answering from generic conventions will miss those details.

A structured bundle does not magically solve that problem.

But it gives the model a better map before it starts reasoning.

Common questions this approach helps answer

A structured codebase bundle is especially useful when you need to ask: what files should I inspect before changing this module, what depends on this file, what does this file depend on, is this module isolated or central, are there circular dependencies around this area, which files appear to be entry points, which modules look like dependency hubs, what directories are most involved in this change, and which parts of the codebase should be included in a given LLM prompt.

These are not final implementation questions. They are orientation questions.

That is exactly why they matter.

A lot of bad LLM coding output starts with poor orientation. The model answers too soon, before it knows where it is in the system.

If you can improve orientation, you can improve the quality of the next question.

Instead of asking an LLM to implement a feature, you can first ask: based on this dependency graph, identify the files that should be inspected before implementing this feature, and explain the dependency path for each one.

That is a better first step. After the relevant files are identified, you can provide source code for those specific files and ask for implementation guidance.

This creates a more disciplined workflow. It also makes the model's reasoning easier to check — if it names a file, it should be able to explain why that file is structurally relevant.

The practical limit of raw code prompts

There is a reason developers keep reaching for bigger prompts: we want the model to have enough information.

But raw code has a poor signal-to-noise ratio for many architectural questions. A single file may contain hundreds of lines of implementation detail that are irrelevant to deciding whether the file matters structurally. Meanwhile, the one import relationship that makes the file important may be easy to overlook.

This is why codebase context should be layered.

Start with structure. Then move to source.

A dependency graph is not the whole answer. It is the first layer of context.

Once you know which modules are relevant, raw source becomes much more useful. The model can read implementation details after it understands why those files are in scope.

That ordering matters.

If you give the model implementation detail before structure, it may over-focus on the visible code. If you give it structure first, it has a better chance of asking the right implementation questions.

That is the workflow this series will build toward.

Takeaway

The useful unit of LLM context is not always the file.

Often, it is the relationship between files.

That is the core reason LLMs struggle with large codebases when the prompt is built from pasted files alone. The model may see code, but it does not automatically see the system around that code. It does not know which files are central, which modules are downstream of a change, which dependencies carry architectural risk, or which paths are worth inspecting before implementation begins.

PViz is built around that gap.

A dependency graph makes those relationships explicit. A PViz bundle packages that graph, along with module and folder-level context, into a structured JSON artifact designed for LLM consumption. Instead of handing the model a pile of source files and asking it to infer the architecture, PViz gives it a repository map first.

That does not eliminate the need to read source code. It makes source code selection more deliberate.

The workflow is not:

paste everything and hope the model finds the important parts.

It is:

use PViz to identify the structurally relevant parts of the codebase, then give the model the source files that matter.

For LLM-assisted development, that is the difference between asking a model to guess and giving it evidence.

Next in the series

In Part 2, we will move from the problem to the workflow.

We will generate a PViz bundle for a real codebase and use it to build a structured prompt. The goal is to show how the bundle changes the starting point for an LLM-assisted code review or implementation planning task.

Instead of asking the model to reason from raw files alone, we will ask it to use the PViz bundle to orient itself first: identify relevant modules, follow dependency relationships, explain why certain files matter, and decide what source code should be inspected next.

That is the practical utility this series will focus on.

PViz is not meant to replace source-level review. It is meant to make source-level review more targeted by giving the LLM a structural map of the repository before the detailed code enters the conversation.

The broader goal is simple: make LLM code analysis less dependent on pasted file dumps and more grounded in the actual structure of the codebase.

Try PViz on your own codebase

Get dependency graphs, coupling signals, and a compressed bundle ready for your LLM — for any GitHub repository, in minutes.