AWS Builder Center

MemoForge: Building a Trustworthy AI Pipeline for Investment Committee Memos

Venture firms need to turn messy founder pitches, notes, and decks into trusted Investment Committee memos—quickly and consistently. But AI prompts can produce polished outputs containing invented numbers, unsupported claims, or formatting errors that silently break workflows. MemoForge solves this by building a verified, self-healing pipeline that checks claims for source traceability, tests plausibility against benchmarks, enforces format compliance, and routes every memo through a human-review QA gate.

Currently an undergrad student pursuing Information Technology with AI/ML honours.

MemoForge: Building a Trustworthy AI Pipeline for Investment Committee Memos

When a venture firm receives a founder’s pitch, the work is rarely just “summarize this deck.” In practice, analysts are asked to turn a messy combination of notes, emails, and slide decks into a clean, house-format Investment Committee memo that a partner can trust quickly. This happens repeatedly, under time pressure, and across analysts of different experience levels.
This sounds like a straightforward automation problem — until you look at the actual failure modes.
A single prompt can produce a memo that looks polished but is built on invented numbers, unsupported claims, or formatting errors that silently break downstream workflows. Worse, many of these problems are invisible at the moment the model generates them. The memo reads well, but the trust signal is false.
This is the problem we set out to solve with MemoForge.

The Problem: AI is Great at Writing, Not Always at Being Trustworthy

Most AI workflows for business writing assume that if the model is fluent and persuasive, the output is usable. In venture or advisory work, that is not enough.
The critical questions are not just:
  • Can the model generate a memo?
  • Can it write well?
The real questions are:
  • Did the founder actually say this?
  • Is this number credible, not just traceable?
  • Is the final output structurally valid and machine-parseable?
  • Is there a human review gate before the memo reaches a partner?
These are very different constraints from “generate something readable.” And they matter because the cost of a bad memo is not just embarrassment — it is bad decision-making.

Why MemoForge Exists

MemoForge is a verified, self-healing, benchmark-aware pipeline for generating Investment Committee memos. It is designed to handle the hardest part of the problem: making AI output not only persuasive, but dependable.
Instead of relying on a single prompt to do everything, we split the workflow into multiple stages, each with a clear responsibility:
  • extract facts from raw pitch material,
  • verify each fact against source quotes,
  • benchmark claims for plausibility,
  • generate a structured memo,
  • repair structural contract violations,
  • enforce a final human review gate.
This architecture reflects a key insight: correctness and credibility are separate concerns.
A fact can be grounded in the founder’s source and still be implausible. A fact can also appear in the deck and still be fabricated by the model. These are different failure modes, and they require different safeguards.

The Design: A Pipeline, Not a Single Prompt

At the heart of MemoForge is a multi-stage agent system.

1. ExtractionAgent

The pipeline begins by converting messy text into structured facts: company details, ask size, market information, operating metrics, product context, and other relevant data points. Each fact is paired with a source quote to preserve traceability.
This matters because if the model writes from raw text without a structured fact layer, it risks producing generic, ungrounded narratives and losing the connection between claims and evidence.

2. VerificationAgent

This component checks whether a claim actually appears in the source material. It asks a narrow but essential question: did the founder say this?
This catches the classic hallucination problem — the model confidently inventing numbers or assertions that are not supported by the pitch.

3. PlausibilityAgent + BenchmarkMemory

This is where the pipeline becomes more interesting.
A claim can be fully traceable to the source and still be unrealistic. For example, a founder may say a startup is growing at 500% annualized growth, and the model might faithfully transcribe it, even though it is far beyond what is credible for the company’s sector and stage.
That is why MemoForge includes a separate plausibility layer that compares grounded claims against benchmark ranges learned over time. The system does not block the claim outright; it flags it for human review. This preserves business judgment while surfacing risk early.

4. MemoWriterAgent

Once facts are verified and flagged appropriately, the memo is written under a strict, machine-parseable contract. That means the output is not just “readable,” but structurally compliant and easy to validate.

5. FormatComplianceAgent

This stage catches contract violations such as missing markers, incorrectly numbered sections, or broken formatting. It is designed to repair issues automatically and quickly, rather than letting broken output pass silently downstream.

6. QACriticAgent

The final gate is intentionally non-bypassable. The system cannot mark the memo as ready without a human review signal. This is a critical safeguard against low-quality, incomplete, or implausible output being mistaken for finished work.
This is not just engineering purity — it is operational discipline. If a memo is going to a partner, it must be not only polished but also reviewable and explainable.

Why This Matters More Than “Better Prompting”

One common approach in AI product work is to add another instruction to the prompt: “check whether this is grounded,” “be careful with numbers,” “be realistic.”
That sounds reasonable, but it is not enough.
Why? Because verification and plausibility are different functions. One answers: “Did the founder say this?” The other answers: “Is this credible in context?” If both are collapsed into a single prompt, they become entangled and weaker. The model may appear to “think” about both, but the mechanism still relies on the same unreliable judgment surface.
MemoForge separates these concerns into independent components. That is one of the project’s most important design decisions.
It reflects a deeper truth: the real challenge in AI-assisted decision support is not just model quality; it is trust architecture.

Evaluating the System

The project was evaluated against a fair baseline: a direct-prompt approach using the same model and same task, but without extraction, verification, plausibility checks, format contracts, or QA gating.
Across a benchmark of synthetic pitch cases, MemoForge outperformed the baseline in ways that matter for real-world usage:
  • every claim in the output was either verified or explicitly flagged,
  • the final output remained machine-parseable,
  • issues were surfaced to a human reviewer instead of hidden,
  • reviewer burden was reduced to a smaller, targeted list of items to check,
  • and the output’s reliability was improved even when the underlying model was not perfect.
This is an important difference between “AI that writes confidently” and “AI that supports decisions responsibly.”

The Honest Tradeoff

One of the strongest things about this project is that it does not hide the cost of reliability.
The system requires more model calls and more orchestration than a single direct prompt. That is an explicit tradeoff: more steps, more structure, and more review burden earlier in the workflow. But the project treats this as a feature, not a flaw.
A partner does not need a memo that looks impressive at the expense of being hard to trust. The team needs a system that can be checked, repaired, and safely reviewed.
That is the real value proposition of MemoForge.

How TrueForge and Qodo Fit In

This project was not built in isolation. It benefited from tools and workflows that support reliable software building and quality engineering.

TrueForge

TrueForge helped reinforce the project’s emphasis on building trustworthy AI workflows, not just generating polished outputs. It fit naturally into the project’s goals by encouraging stronger validation, higher confidence in system behavior, and a more disciplined approach to model-assisted reasoning.
In other words, it aligned with the central principle behind MemoForge: AI should be used as a tool for structured reasoning under supervision, not as an unchecked writer.

Qodo

Qodo helped improve the quality of the implementation itself. As the project expanded across multiple agents and validation stages, code quality became a real challenge. Qodo supported safer refactoring, better logic review, and more robust engineering practices across the pipeline.
This was particularly important because the project relies on multiple moving parts: extraction logic, benchmark checks, contract validation, repair loops, and QA logic. Quality tooling helped keep that complexity manageable.

Lessons Learned

The biggest lesson from building MemoForge is that AI systems for decision support need architecture, not just prompting.
A prompt alone is not a workflow. A workflow needs:
  • explicit structure,
  • verifiable evidence,
  • benchmark-aware risk checks,
  • output contracts,
  • and clear review gates.
We also learned that “generated output” is not equal to “usable output.” In a high-leverage environment like venture analysis, the value lies not in producing attractive language, but in producing language that can be trusted, checked, and acted on.
The project’s most important insight is this:
Trust in AI workflows is created by layered systems — not by hoping the model behaves well.

Looking Ahead

MemoForge is a strong foundation for building more trustworthy decision-support systems. The next step is not simply “make it more creative.” The next step is to make it more robust, more domain-aware, and more operationally realistic.
That could include:
  • richer benchmark memory for more sectors and stages,
  • better source grounding and metadata extraction,
  • more nuanced human-review routing,
  • and broader application beyond venture memos to other evidence-heavy reporting workflows.
The real opportunity is not to replace analysts. It is to help them move faster without giving up judgment, reliability, or traceability.

Final Thoughts

MemoForge started with a practical pain point: the cost of turning messy founder material into a credible IC memo is high, repetitive, and vulnerable to silent AI failure. The project addresses that challenge by creating a pipeline designed around evidence, structure, and review.
It is not a “chatbot for VC memos.” It is a trust-aware writing system for professional decision support.
And that is the difference between AI that sounds useful and AI that is actually usable.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article