← Back to research

Research systems / DEEPSIGHT RESEARCH

The scarce resource for research agents is reliable feedback

Candidate text and synthetic data can expand quickly. The scarce part is defining quality, finding long-tail failures, and connecting corrections to real research work.

The judgment

“We have a large volume of high-quality data” is one of the most common and least testable claims in AI.

Quality is not a permanent label attached to a dataset. It depends on the objective, task, date, validation method, and cost of error. A system can generate unlimited candidate answers. The candidates cannot decide for themselves what is true, material, or ready to enter a conclusion.

For a research agent, the scarce resource is reliable feedback:

  • Who defines whether the work is good?
  • Which errors can tools detect?
  • Which judgments require expert review?
  • Does the feedback cover long-tail risk?
  • Can a correction improve the next assignment?

Synthetic data expands the candidate supply

Models can generate questions, answers, research plans, counterarguments, and report drafts. This can increase training and evaluation coverage and reduce the cost of basic labeling.

The limitation is fundamental: candidate generation does not create an external ground truth.

When the verifier is dependable, synthetic data can help search for better candidates. When the verifier rewards fluency, formatting, or surface agreement, the system can scale the wrong behavior faster.

Synthetic data is best treated as a candidate generator, not as its own factual authority.

Feedback exists at four levels

1. Formal feedback

Check whether required files exist, formats are valid, citations are present, and calculations execute. This is the easiest layer to automate.

2. Factual feedback

Check whether a source supports a claim and whether entity, period, currency, and denominator are consistent. Retrieval and deterministic tools need to work together.

3. Judgment feedback

Check whether the evidence justifies the business interpretation, whether counterexamples are missing, and whether an inference crosses its factual boundary. This usually requires a domain rubric and expert review.

4. Outcome feedback

Observe whether the research improved diligence, investment committee preparation, risk identification, or an operating decision. This is the most valuable feedback and also the slowest and most confounded.

Optimizing the first two layers can remove many basic errors. Professional capability requires some route for the latter two to influence the system.

Expert labels are contextual

Experts disagree. Investment strategy, risk tolerance, decision rights, and research stage vary across institutions. A conclusion that helps a growth investor may not fit a strategic investor. A report suitable for initial screening may fail formal diligence.

A reliable feedback system should preserve:

  • who provided the feedback;
  • the task and stage in which it occurred;
  • whether the correction concerned fact, inference, or presentation;
  • the institution and time context;
  • unresolved disagreement.

Collapsing every expert revision into a universal correct answer removes exactly the conditions that make professional feedback useful.

The expert's role moves upward

Constitutional AI is one example of moving some human feedback from repeated case-by-case correction into principles and evaluation procedures.

In an investment research system, experts can similarly move from editing every sentence toward:

  • defining source and evidence qualification;
  • designing report and judgment rubrics;
  • reviewing conflicts and long-tail errors;
  • deciding which corrections should become reusable rules;
  • handling exceptions that validators cannot cover;
  • deciding what may be delivered without further review.

This does not remove the person. It places expert time where it changes the result.

Feedback becomes an institutional asset only after reuse

One correction helps the next assignment only when it is recorded, attributed, and tested again.

Suppose an analyst finds that companies in a sector routinely blur contracted value and recognized revenue. If the correction remains only in final prose, the next assignment can repeat the error. If the system preserves the error type, applicable context, check, and counterexample, it can become a quality gate.

The useful loop is:

real assignment → candidate research → expert correction → error attribution → rule or evaluation update → validation on a new assignment

Without attribution and new-task validation, a “data flywheel” may be only a growing pile of text.

Boundary

This note does not assume that every expert correction should become a rule. Reusable research discipline and context-dependent investment judgment must remain separate. The test of a feedback system is not how much it remembers, but whether it improves a new assignment without spreading an old mistake.

Research note

DeepSight separates external fact, inference, judgment, and open questions. Links in the article lead to cited source material. This note is not investment advice.