Claude

IntuitionLabs is now a member of the Claude Partner Network – AI training and upskilling with Claude for pharma and biotech. Book a call.

IntuitionLabs
8/7/202620 min read

Two Tools, One Workflow: Using Claude Cowork and ChatGPT Together

ai workflowsclaude coworkchatgptai governanceai verificationcommercial operationspharma ai adoptionhipaa compliance

Executive Summary

This piece makes one argument: pairing two AI tools deliberately, one to produce, a different one to independently verify, catches errors that trusting a single tool never would. The demonstration below is a real one, not a simulation. Claude Cowork drafts a one page briefing memo from a Deloitte report on its own: opening the source file, writing the document, saving it to the correct folder without being told twice. ChatGPT then checks two of that memo's central claims independently, and finds real problems in both. What follows covers why that pairing beats single tool reliance, why Claude Cowork and ChatGPT specifically for a pharma commercial operations audience, the walkthrough itself, what happens when nobody runs that verification step (four real, documented cases), what is actually happening inside Cowork during each step of the demonstration, and where else this pattern fits for this same audience, including where compliance draws a hard line around it.

Why Ensemble: Matching Tools to Task, Not One Tool for Everything

Commercial Ops teams evaluating AI vendors are not choosing a single tool to do everything. The stronger pattern, and the one this walkthrough demonstrates, is matching each tool to what it does best, then using a second model to check the first. Here, Claude Cowork handles autonomous, file native production: reading a source document, drafting, saving to the right folder. ChatGPT handles fast, independent checks on individual claims to catch anything Cowork got wrong before it reaches a reader.

The assumption worth naming, and challenging, is that one AI tool ought to handle everything a team needs. That assumption doesn't hold up. Practitioners running large model deployments in production have started calling this "the 19 model problem": no single model wins every task category once cost, speed, structured output, and security boundaries are all weighed together, and enterprises are increasingly routing work across models rather than standardizing on one ([1]). Gartner's own forecasting points the same direction: the firm predicts 40% of generative AI solutions will be multimodal by 2027, and separately notes that narrower, domain-specific models carry lower hallucination risk than general purpose ones asked to do everything ([2]). McKinsey's framing of multimodal AI makes a related point from a different angle: combining sources gives a more complete grounding for a claim than any single source alone, including a healthcare example where fusing modalities measurably outperformed any single modality on its own ([3]).

That's the general case for matching tools to tasks. The workflow demonstrated here makes a narrower, more specific claim: not just that different tools suit different tasks, but that having a genuinely independent second model check a first model's output catches errors that checking your own work never will. This is closer to an active area of AI research than a hunch. A 2025 paper on cross-model consistency, built specifically around comparing outputs across different models rather than having one model check itself, reports a 6 to 39 percent improvement in hallucination detection accuracy over prior single-model detection methods ("Finch-Zk," arXiv 2508.14314). The underlying logic is straightforward: two models trained differently are less likely to share the same blind spot, so where they disagree is itself a signal worth investigating. A related and frequently cited method, "LM vs LM: Detecting Factual Errors via Cross Examination" ([4]; full PDF), has one model adversarially interrogate another's claims to surface factual inconsistencies between them, the same basic shape as handing a memo's headline number to a second model and asking it to check the math.

It's worth noting that even a single model checking its own work measurably helps. Research on "chain of verification" shows a model that drafts an answer, generates its own independent verification questions, answers those separately, then revises its original answer, produces fewer hallucinations than one that doesn't run that self-check ([5]). If self-verification alone helps, a genuinely separate model, with no shared blind spots to protect, should help more. One 2025 framework that cross-checks model output against external databases, live search, and academic literature in real time reports a 67 percent reduction in hallucinations without any loss in response quality ([6]); that figure comes from a single paper rather than a settled consensus number, so treat it as illustrative of the mechanism's potential rather than a guarantee, but the direction of the evidence is consistent across all of it.

None of this requires an enterprise AI strategy to act on. One practitioner described the pattern in plain terms after building the habit without planning to: ChatGPT for structured reasoning and first drafts, Claude for long form writing and tone, a deliberate two tool split rather than a single default ([7]). The workflow below follows the same instinct, with the verification step made explicit and deliberate rather than incidental.

Why Claude Cowork and ChatGPT, Specifically

Worth being upfront about what isn't known: there is no public, brand-named usage data showing how pharma or biotech commercial operations teams specifically split their work between Claude and ChatGPT. That granularity doesn't exist yet in credible public sources, and claiming otherwise would overstate the evidence. What does exist is a clear picture of the stakes and the gap this workflow is built to close.

McKinsey's survey of more than 100 pharma and medtech leaders found that 100 percent had experimented with generative AI by late summer 2024, 32 percent had scaled a use case, and only 5 percent reported the technology as a realized competitive differentiator with consistent, measurable financial value ([8]). That gap between experimentation and realized value is the actual problem this article is addressing: adoption was never the bottleneck, workflow design is. McKinsey separately estimates that 18 to 30 billion dollars a year of the sector's generative AI value potential sits specifically in commercial functions, this article's stated audience, not in R&D or manufacturing ([9]). Deloitte's own research, a different report from the one this walkthrough's memo is built on, estimates a large biopharma company could realize up to 7 billion dollars in value over five years through generative and agentic AI, with 20 to 30 percent cost savings in targeted functions ([10]), and has written specifically about whether AI can make biopharma sales reps more effective, direct territory for a commercial ops and field force audience ([11]).

A wider, more sobering data point belongs here too. A widely cited 2025 analysis of enterprise generative AI deployments found that roughly 95 percent of pilots showed no measurable profit and loss impact, against tens of billions of dollars in enterprise spend. Read alongside McKinsey's 5 percent realized-differentiator figure above, the pattern is consistent: the tool isn't what separates the pilots that work from the ones that don't, the workflow built around it is.

That's the case for why this matters. The case for these two tools specifically comes down to role fit, not brand preference. Cowork's value in this workflow is autonomous, file native production, opening a source document, drafting, and saving to the right place without a human doing the mechanical work. ChatGPT's value is fast, independent verification with web search available, checking a specific claim against outside sources without any stake in having drafted it. Readers who want the fuller platform-by-platform comparison can find it in IL's own Claude vs ChatGPT vs Copilot vs Gemini: 2026 Enterprise Guide and ChatGPT Enterprise vs Claude Enterprise: Feature Matrix; this piece isn't re-deriving that comparison, it's demonstrating one specific, deliberate way to use both together. For scale, ChatGPT alone has more than 1 million business customers and roughly 900 million weekly users as of early 2026. And this isn't a hypothetical for life sciences specifically: more than half of life sciences staff surveyed in 2025 reported using ChatGPT at least monthly, often despite organizational restrictions on doing so. Informal, ungoverned use of exactly this kind is already happening; the workflow below is one way to make it deliberate instead.

Step 1: Start from a real source document

The task: turn a downloaded report into a one page briefing memo, flagging the two or three findings most relevant to AI adoption for a specialty pharma, leaner field force context. The source was a Deloitte report, "Accelerating the Future: End-to-End Transformation of Pharma's Commercial Activities" (Life Sciences and Healthcare Predictions 2030), 2024.

Cowork prompt: read the Deloitte PDF report and turn it into a one-pager briefing memo

Step 2: Let Cowork do the production work

Rather than hand writing a Word file, Cowork calls its packaged docx skill, a purpose built tool for document generation, and this shows up directly in its tool call log. It opens the Deloitte PDF itself and runs its own file handling commands before writing anything, then checks its own output ("Only one page, good") before calling the draft finished. That self check happens without a request for approval, so nothing here has been reviewed by a person yet.

Cowork reading the source PDF before drafting

Cowork calling its packaged docx skill to produce the memo

Cowork's self-check ("Only one page, good") before calling the draft finished

A detail worth calling out: the original prompt named a specific folder inside the shared repository, not just "the file." A repository can carry standing designations, for example which folder holds drafts pending review versus which folder holds confirmed, team facing artifacts. Set that once, through memory or written instruction, and Cowork respects it on every run without being told again. That is why this memo lands in the drafts folder rather than the team shared one.

Step 3: Review the first draft

The result: one page, addressed to a Commercial Ops leader, three findings pulled from the Deloitte report. Cowork saves it to the drafts folder automatically, pending review, not yet a team facing artifact. This is the version that goes to fact checking next; nothing in it has touched a second model yet.

The finished one-page draft memo, saved to the drafts folder pending review

Step 4: Pivot to ChatGPT to verify the headline number

The memo's central figure, that a top ten biopharma company could capture $5 to 7 billion in peak value over five years, with 35% from the commercial function, gets pulled out and handed to ChatGPT with web search on.

ChatGPT web search verifying the memo's headline $5-7 billion figure against Deloitte's own material

Finding: directionally correct, but imprecise. ChatGPT confirms the underlying Deloitte figures, then corrects the framing: 35% is the commercial function's share of the total value opportunity, not a separately derived commercial only estimate. A subtle overstatement, caught before it reached the reader.

Step 5: Check a second claim the same way

One claim is not a habit, so the workflow repeats the check on the memo's peer ambition figures: $200 to 500 million in annual content generation savings, revenue leakage recovery, and 10 to 25% faster launch timelines.

ChatGPT's claim-by-claim verification table for the memo's peer-ambition figures

Finding: partially supported, not fully. ChatGPT's claim by claim table finds the content savings and launch timeline figures accurate, but flags "recovering billions" in rebate payments as stronger language than Deloitte actually used. Deloitte describes an optimization opportunity, not an achieved recovery. Two rounds, two real corrections; this is verification catching something, not verification as a formality.

Step 6: Send the corrections back into the memo

Cowork rewrites both findings against ChatGPT's feedback. Finding 1 now reads as a modelled value opportunity, explicit that it is projected, not observed. Finding 2 now says Deloitte "cites peer ambitions, not proven outcomes," matching Deloitte's actual wording on rebate payment optimization. The memo stays one page, and still saves to the drafts folder under the same designation as before; a correction does not automatically promote it to team shared.

Cowork confirming both findings were rewritten and re-saved to the same drafts folder

Step 4 and Step 5 above are the part of this workflow that actually matters. ChatGPT caught two real overstatements in a memo Cowork had already marked finished, not typos or formatting issues, but the kind of subtle framing errors that read as correct on a first pass. What follows is what happens when nobody runs that second check.

When Nobody Checks: What Ungoverned AI Use Actually Costs

Samsung banned ChatGPT company-wide in 2023 after employees pasted proprietary source code, a transcribed confidential meeting, and internal test sequence data into the tool across three separate incidents in twenty days ([12], CIO Dive). The incident is formally documented in the AI Incident Database ([13]). This is the failure mode on the input side: nothing was wrong with what came out of the tool, the problem was what went into it, with no policy in place to stop it.

The output side has its own well documented failure. In Mata v. Avianca, two attorneys submitted a federal court brief containing six entirely fabricated case citations generated by ChatGPT. A judge in the Southern District of New York sanctioned both lawyers and fined them 5,000 dollars ([14]; legal industry summary via Seyfarth Shaw). Nobody checked the output against a source before it went to a federal judge. It's the mirror image of what Step 4 and Step 5 did right in the walkthrough above.

It isn't only individuals who skip this step. In 2025, Deloitte was forced to refund part of a roughly 290,000 dollar report it had produced for the Australian government after the report was found to contain fabricated academic citations and a fake court quote, generated using Azure OpenAI's GPT-4o. Deloitte disclosed its AI use in a revised version of the report ([15]; formally recorded in the OECD.AI Incidents database). It's worth naming this one directly rather than talking around it, including because of the obvious tension it raises: this article is itself an AI drafted, AI checked piece, discussing a firm that skipped exactly the check this workflow is built around. That tension is the point, not an embarrassment to avoid. A major professional services firm, with every resource available to catch this before publication, didn't run an independent verification pass, and it cost them a client refund and public correction. The two-tool workflow above exists specifically to prevent that outcome, and this article's own Step 4 and Step 5 are that prevention in action, not just a claim about it.

The last case ties the first three together. In 2024, a Canadian tribunal held Air Canada legally liable for its own chatbot's fabricated bereavement fare policy, ruling that a company is responsible for what its AI tells a customer regardless of which human or system produced the error ([16]; legal framing via the American Bar Association). Whether the failure is a leak, a fabrication, or an unchecked professional deliverable, the liability lands on the organization that shipped it, not on the tool.

None of this is a fringe concern inside enterprises. As of early 2023, only 3 percent of firms had banned ChatGPT outright, and roughly half were still drafting a policy for it at all ([17]); a 2024 Forrester and Grammarly survey found 32 percent of organizations cite security concerns and 27 percent cite a lack of policy as active blockers to AI adoption, the same source notes. Readers building a policy rather than just reacting to one incident at a time can start with IL's own Developing a Corporate AI Policy and Governance Guide and AI Hallucinations in Business: Causes and Prevention.

Inside Cowork: What's Actually Happening in Steps 2 Through 6

Now that you've seen it work, it's worth explaining what was actually running underneath each of those moments.

The folder detail in Step 2 wasn't incidental. The original prompt named a specific folder inside the shared repository, not just "the file," and Cowork respected that designation without being told twice on a later run. That's memory and standing instruction at work: a repository can carry standing designations, for example which folder holds drafts pending review versus which folder holds confirmed, team facing work, and setting that once means Cowork applies it going forward rather than needing the instruction restated every time.

The file handling in the same step is the other half of what makes this "file native" rather than chat only. Cowork opened the Deloitte PDF itself and ran its own file handling commands before writing anything, the same category of autonomous action that let it save the finished memo to the correct folder without a human moving the file afterward. This is the structural difference from a tool that only produces text in a chat window and leaves the file handling to a person.

The docx skill mentioned in Step 2, a packaged, purpose built tool for document generation that shows up directly in Cowork's own tool call log, is an example of a broader pattern: specialized, callable skills as an extensibility layer, rather than one general purpose chat interface trying to do document generation, spreadsheet work, and file management all through the same undifferentiated conversation.

And the self-check in Step 2, Cowork confirming to itself "only one page, good" before calling the draft finished, is worth reading against the research on self-verification described earlier in this piece. That kind of same-model self-check is real and it measurably helps; the chain-of-verification research cited above shows a model that checks its own answer produces fewer errors than one that doesn't. But it is not the same thing as what happened in Step 4 and Step 5, where a genuinely separate model, with no stake in having drafted the memo, checked it. Cowork's self-check happened without a request for approval, so nothing in that memo had been reviewed by a person, or by a different model, at that point. That distinction, self-check versus independent check, is the entire argument of this article compressed into one moment in the walkthrough.

Beyond This Memo: Where Else This Pattern Fits, and Where It Doesn't

The memo above is one use case. The same draft-then-verify pattern extends naturally to other work this audience already does. HCP targeting and next-best-action models increasingly adjust channel mix on a rolling basis rather than an annual refresh, exactly the kind of output worth having a second tool check against the underlying data before it drives a real budget decision ([18]). Pre-call planning and incentive tracking are another plausible fit: draft a rep's pre-call brief with one tool, verify its claims against actual CRM data with a second before it reaches the field ([19]).

This pattern isn't hypothetical inside this ICP's own CRM stack, either. Veeva launched a Quick Check Agent and a Content Agent for Vault PromoMats in December 2025, a pre-review agent that scans draft content for editorial, branding, and regulatory issues ahead of formal MLR review, paired with a second agent that provides context-aware review and summarization ([20]; fuller detail in IL's own Veeva AI Agents: Agentic AI for the Life Sciences Industry and Automating MLR Review with Veeva PromoMats AI Agents). Moderna's Global Marketing Operations Director, Jason Benagh, described the early access experience directly: "What we've learned as an early access user is that the Veeva AI Quick Check Agent moves Moderna closer to a process where parts of MLR could become nearly touch-free... now we can genuinely see how we might get there." Worth being precise about one detail here: Veeva's AI agents are LLM-agnostic, run on AWS Bedrock or Azure depending on customer choice, not a confirmed Anthropic deployment. The relevance is structural, not a shared vendor: a staged, draft-then-verify pattern is already live in this industry's own CRM infrastructure, independent of which specific tools this article covers.

That structural similarity has a hard boundary, and it's worth being exact about where it sits. The FDA gives AI generated content no separate compliance lane: the same fair balance and claim support rules apply regardless of what drafted the content, and no regulator treats AI as a final approver. FDA enforcement activity around promotional claims hit its highest level in nearly 25 years in 2025, with more than 200 enforcement letters issued ([21]). A second AI tool checking a first tool's claims is a real improvement over no check at all, but it is not, and cannot be, a substitute for the human MLR sign-off this industry's own compliance process requires.

The same precision applies to the HIPAA caveat already raised about this specific workflow. Anthropic's own support documentation is direct on this point: Cowork is not an eligible service under Anthropic's Business Associate Agreement in any configuration, regardless of retention settings or which model is in use, and shouldn't be used with protected health information ([22]). That is a genuine, current limitation, not a documentation gap. Anthropic does offer BAA coverage elsewhere, but only through specific, separately activated surfaces: HIPAA-ready Claude Enterprise plans and the first party API, once an organization's primary owner has signed the BAA and turned on HIPAA compliance in settings ([23]; Morgan Lewis, "Healthcare AI Deployment"). Cowork and a BAA-activated Enterprise plan are different products with different compliance postures under the same company; one covers PHI under contract, the other explicitly does not, in every configuration. This workflow used a public, vendor neutral Deloitte report, not patient, prescriber, or other regulated data, deliberately, for exactly this reason. Treat this as current as of today, not a permanent architectural fact: compliance scope for AI tools moves quickly, and it is worth re-checking directly against Anthropic's documentation before every new use case rather than assuming from a prior run. Readers building out a fuller compliance architecture around any of this can start with IL's own Claude for Healthcare and Life Sciences: 2026 Technical Guide and Private LLM Pharma Compliance Architecture.

Same discipline as the workflow itself: verify before you rely on it.

External Sources (23)
Adrien Laurent

Need Expert Guidance on This Topic?

Talk to IntuitionLabs about how to put this guide into practice on your team.

I'm Adrien Laurent, Founder & CEO of IntuitionLabs. With 25+ years of experience in enterprise software development, I specialize in creating custom AI solutions for the pharmaceutical and life science industries.

DISCLAIMER

The information contained in this document is provided for educational and informational purposes only. We make no representations or warranties of any kind, express or implied, about the completeness, accuracy, reliability, suitability, or availability of the information contained herein. Any reliance you place on such information is strictly at your own risk. This document may contain content generated with the assistance of artificial intelligence technologies. AI-generated content may contain errors, omissions, or inaccuracies. Readers are advised to independently verify any critical information before acting upon it. All product names, logos, brands, trademarks, and registered trademarks mentioned in this document are the property of their respective owners.

© 2026 IntuitionLabs. All rights reserved.