Running research Shipped
What you type, what it costs, and what comes back.
myteams researchteam research "loop engineering for long-running coding agents"
myteams researchteam findings
myteams researchteam insights --specialist verification-researcherWhat comes back
A relevance-ranked report of claims, each with its source and a validated confidence, plus an explicit open-questions section. Everything in the Research root of this site came out of runs like this — including the pack that failed, which is still published because a failed research run with a stated cause is more useful than no record of it.
The evidence bar is domain-appropriate
A claim about a published algorithm and a claim about what a vendor shipped last week do not deserve the same standard, and the team does not apply one. What it does apply uniformly is that a claim without a source is not a claim. The pack whose web access was denied mid-run is the proof: all 119 of its trust-scored claims shipped carrying the denial in the text — arXiv ids marked as recalled and unconfirmed, quotes marked as approximate recollection rather than captured verbatim, two citations flagged outright as fabrication risks to verify or drop. That is a different artifact from one that quietly presents recall as research.
Pinned to Claude, for now
The team pins its backend explicitly rather than inheriting the global default, with a note in its own definition saying to remove the pin once the alternatives are hardened. Research quality is the thing most sensitive to model capability, so it gets the least experimentation.