Toward Governed Human–AI Cooperation in Science: Why the Memorandum's Framework Is Needed Now

Two parallel LinkedIn discussions this week — one from Dr. Shekeeb Mohammad, a Paediatric Neurologist and Associate Professor, on AI in peer review, and another from Shahnwaz Ali, Director of Product at Macquarie Bank, on Agentic AI more broadly — converge on the same underlying truth from opposite directions: the era of treating AI as either a forbidden shortcut or a full replacement for expert judgment is over. What remains open, and what neither thread fully resolves on its own, is how to build that cooperation correctly — which is precisely the discussion, development, and implementation our memorandum calls for.

Dr. Mohammad's Experiment: A Real Finding, Wrongly Generalized

Dr. Mohammad tested AI as an unsupervised, standalone peer reviewer, feeding it manuscripts he had already reviewed himself along with his own prior reviews to learn from. His finding was clear: AI missed core issues, misread authors' aims, and produced generic, template-like commentary that could apply to almost any observational study. His conclusion — "DO NOT RELY on AI for peer review" — is a valid description of that specific failure mode, but it overreaches into a blanket prohibition. His confidentiality objection stands entirely on its own merits and is not in question. But the experiment itself only tested one configuration: AI as full decision-maker, asked to do the judging, not merely the checking. That configuration failing tells us little about a different, untested configuration — AI as a literature-connected assistant supporting a human reviewer who retains final authority.

What the Agentic AI Discussion Adds

Shahnwaz Ali's post argued that rising AI autonomy makes deep domain expertise more valuable, since AI can plan and execute but only expertise can judge whether the outcome is right. The far more useful contribution, though, came from the replies. Greg Monahan offered a sharper correction: "The more autonomy we give to AI Agents, the more controls we should put in place... Domain expertise remains as valuable as it was, it doesn't become more valuable. An agent can possess domain expertise. Thus agents can suggest guardrails and review outcomes... Ultimate responsibility lies with the human in charge, of course." This reframing matters: the point isn't that expertise magically appreciates — it's that autonomy without proportional control is the actual danger, and accountability must stay pinned to an identifiable human no matter how capable the tool becomes. Ellis Ruffley added the key distinguishing skill: "The roles gaining value are the ones where judgement is the product: people who can tell whether an agent's output is actually right, not just plausible." That single distinction — plausible versus actually right — is exactly the failure Dr. Mohammad documented. AI's peer-review commentary looked reasonable on the surface but lacked the grounding to be correct.

Putting the Two Threads Together

Read side by side, these discussions aren't in conflict — they're describing the same transition from different vantage points. Execution is being commoditized across every knowledge field, including scientific review. What isn't being commoditized, and what is becoming scarcer and more valuable, is the ability to verify that AI-generated output is actually correct, to catch where it's subtly wrong, and to be accountable for the final call. Dr. Mohammad's experiment doesn't prove AI has no place in review — it proves that AI cannot occupy the judgment seat. It says nothing about whether AI, restricted to the mechanical work of database cross-referencing, citation verification, and inconsistency-flagging, could make the human reviewer faster and more thorough while the human alone decides what the findings mean.

Why This Points Directly Back to the Memorandum

Both threads are groping toward a structure they never quite name, and that structure already exists in our memorandum, Memorandum on Personalized Responsibility, Unified Oversight, and Human–Artificial Participation in the Global Scientific Literature Space — itself produced with AI assistance, a small demonstration of the principle it argues for. Its core standards map directly onto what both discussions are missing:

  • AI participation must always be disclosed, never presented as if it were the reviewer's own unaided judgment.

  • Responsibility for any review, and its consequences, remains with the identifiable human reviewer — never transferred to the tool, regardless of how autonomous or capable it becomes.

  • Review itself is scientific work with its own standing, meaning the human's synthesis and judgment, not the AI's draft, is the actual intellectual contribution being evaluated.

  • Confidentiality and data protection rules must be honored regardless of which tools are used — non-negotiable, independent of how good AI review eventually becomes.

These are not abstract principles waiting for some future need. Dr. Mohammad's failed experiment and the Agentic AI thread's unresolved debate about where judgment lives are live evidence that the profession needs exactly this kind of governed framework now, not later. Isolated experiments and LinkedIn opinions, however well-intentioned, cannot substitute for the sustained discussion, refinement, and institutional adoption a framework like this requires.

The Correct Way Forward

The right conclusion isn't "avoid AI" or "let AI decide." It's that the profession now needs the discipline both threads are circling: build workflows where AI handles volume and verification, humans handle judgment and accountability, and every step in between is disclosed and governed rather than improvised. That is the cooperation the moment calls for — and building it correctly is exactly why the memorandum's principles need active discussion, development, and implementation across journals, institutions, and the broader scientific community, rather than remaining a proposal on the shelf while individual experiments and opinions fill the vacuum in the meantime.

An Open Invitation

The memorandum was deliberately written not as a finished doctrine but as an open call for reflection and cooperation, and that invitation stands here as well. Scientists, reviewers, editors, publishers, institutions, and AI developers who engage with these questions in their own daily practice — as Dr. Mohammad and the Agentic AI thread's contributors are already doing, even without naming the framework — are the exact constituency this proposal needs in dialogue. The memorandum's authors welcome critique, counter-examples, and refinement from the broader scientific community; its value will ultimately be measured not by how completely it anticipates every case, but by how much genuine discussion, testing, and revision it provokes as human–AI cooperation in science moves from informal experimentation toward accountable, governed practice.


Comments

Popular posts from this blog

Menopause Is Not a Gumboil: Answering Clinical Misunderstandings in Light of Medscape

The Excellence of My Age

Two Sides of Frailty: Vulnerability, Compensation, and a Consciousness-Centered Medicine of Aging