What this is, in one minute
A Singapore SME owner has a few recurring chores: work out cash flow, reply to customers, and, if GST-registered, file the quarterly GST F5 tax return. AI tools promise to help. I put four realistic setups head-to-head on the same tasks, using 5 made-up (but realistic) Singapore businesses: a bubble-tea shop, a minimart, an aircon trades firm, a design studio, and a bookkeeping back office.
I ran ~280 graded runs across four rounds (quality, GST safety, cash-flow/drafting safety, and a live connector probe, see "What was tested" below), then re-ran the safety rounds in July 2026 on a newer model to see what had changed. Every answer was produced by Claude Sonnet and independently graded by Claude Opus. All the business data is synthetic and Singapore-localised (SGD, GST 9%, PayNow, WhatsApp). No real company data. The headline moved between the two runs, and that shift is itself the main finding (③).
How to read this: ① context (here) → ② the framework (three layers) → ③ the evidence (the runs) → ④ tool-by-tool (what an AI connector can actually reach).
The four setups compared
What was tested: every combination
| Round | Question | Setups compared | Workflows | Runs |
|---|---|---|---|---|
| 1 · Quality | Which gives the best output? (expert prompt) | plain · generic skills · plugin · plugin+script | cash-flow · drafting · GST | 28 |
| 2 · GST safety | How often would it misfile GST? (naive owner) | plain · plugin · plugin+script · plain+script | GST | 160 (80 free-form + 80 primed) |
| 3 · Safety beyond GST | Does the guardrail help on non-tax work? | plain · plugin · plain+script | cash-flow · drafting | 90 |
| 4 · Live connector | Does the Xero MCP inherit SG config? | + MCP (Xero) | live org probe | SG + Global orgs |
The per-setup coverage grid (what's measured vs still reasoned) is at the foot of ③ What I found.
Plain-language glossary
bottom line It's not one tool, it's a combination, and it moves over time
Keep the status quo (Xero) for the books & filing: only it produces a deterministic, audit-trailed number and files it to IRAS. Add plain Claude for the thinking: cash-flow questions, customer replies, what-ifs. Add a thin SG guardrail for anything customer- or forecast-facing: on the mid-2026 re-run Claude got GST right unguarded (the old catastrophic-underpay finding no longer reproduces, a knowledge gap the model closed by itself), yet it still fabricated forecast confidence bands and over-promised refunds to customers in nearly every run, behavioural flaws a thin guardrail eliminated. Skip the small-business plugin for the books: its 10 finance skills carry zero SG tax logic (I read them), it adds friction, and its US-shaped format was the worst offender on invented forecast precision. The pattern to remember: knowledge gaps close with every model release; behavioural and structural gaps don't.
The comparison: three layers, not one tool
The clearest way to see the difference: a setup is built from three independent layers. Most "tools" only change one layer. Below, first the building blocks (the intuition), then the detailed matrix (the verdict per front). The run-by-run rates behind these verdicts live in ③ What I found.
A. The building blocks
Claude's intelligence is the constant, and it improves over time (the mid-2026 re-run in ③ shows the tax-knowledge gap closing on its own). Everything else is a choice in one of three layers. An MCP connector upgrades the Data layer; the small-business plugin's US format downgrades Processing (fabricated precision + friction); an SG guardrail upgrades Processing (bans invented precision & over-commitment). No AI option clears the Filing gate.
Each setup = one pick per layer
your question "What if we run the plugin with MCP? Doesn't account config fix it?"
The data gets fixed. The format doesn't. An MCP on an SG Xero org pulls SG data via account config (SGD, GST 9%) even under the plugin. At the current model Claude also handles the SG tax correctly under the plugin: the old "US process overrides SG knowledge" harm has faded (see the re-run in ③).
But the plugin's output format is the live problem: it fabricates forecast confidence bands in every cash-flow run (4/4, see ③), now dressed on authoritative-looking live data. The fix isn't a connector, it's the Processing block: drop the plugin, use Claude's own knowledge + a thin SG guardrail that bans invented precision and unauthorised commitments. That you can change; the missing local connectors and the CorpPass filing gate you can't.
B. The detailed matrix
| Layer / front ↓ | ① Plain Claude m | ② + plugin skip m(± MCP) | ③ + MCP re.g. Xero connector | ★ + SG script best Claude m(measured) | ④ Status quo keep vXero + accountant |
|---|---|---|---|---|---|
| ① Data: where the numbers come from | |||||
| Auto-pull vs paste | ✗ paste CSV by hand | ~ live only via its US connectors 📦 | ✓ live auto-pull | ~ paste, or add MCP | ✓ native |
| SG data (account config) | ✓ you paste SG data | ✓ data is SG (process isn't) | ✓ inherits SG org config | ✓ SG data | ✓ native SG |
| Reach SG-local tools PayNow·WhatsApp·IRAS | ✗ manual export | ✗ no such connector | ✗ no connector exists | ✗ can't be scripted in | ✓ used natively in the apps |
| ② Processing: how GST is handled | |||||
| SG tax logic | ✓ Claude knows it (confirmed on re-run) | ✓ model overrides the US text | ✓ Claude knows it | ✓✓ reinforced by guardrail | ✓ native |
| GST safety won't underpay IRAS | ✓ now correct unguarded (re-run 0/7 catastrophic) | ✓ now correct too (0/6) | ✓ correct on a clean pull | ✓✓ correct + guardrailed | ✓ deterministic |
| Behavioural safety no invented precision / over-commit | ✗ over-commits 4/4 on replies | ⚠ worst: bands 4/4 + over-commits 4/4 | ✗ same model behaviour (no guardrail) | ✓ 0/4 both, guardrail eliminates | ✓ deterministic outputs |
| Numeric accuracy revenue sum | ✓ revenue exact w/ tools (~100%) | ✓ revenue exact too | ✓ reliable on a clean pull | ✓ revenue exact (10/10) | ✓ guaranteed + audit-trailed |
| Analysis & drafting | ✓✓ best in eval | ✗ fabricates forecast precision + friction | ✓✓ best + live data | ✓✓ best + guardrails | ✗ rigid (JAX/Assist emerging) |
| ③ Filing: submit to IRAS | |||||
| File GST F5 → IRAS | ✗ can't file | ✗ can't file | ✗ not in tool's reach | ✗ can't file | ✓ Xero one-click (ASR+) |
| Overall | |||||
| Setup & cost SGD/mo | ✓ ~zero · ≈S$0–28 | ~ install + connectors | ~ OAuth per tool | ✓ paste a recipe · ~zero | • S$39–95 + accountant · grant support may apply |
| Best role | Thinking layer | Skip: downgrades process | Thinking + live reads | Safest Claude config for SG tax | Books & filing: keep |
The "★ + SG script" column is directly measured (plain Claude + the SG script, no plugin, see the safety benchmark on the results tab); it matches plugin+script, confirming the plugin adds nothing. The "+ MCP" connector/data layer is now confirmed live (a Singapore Xero org returned SGD + GST 9% from config; 68 invoices pulled live; no filing endpoint, see results tab). Only the reasoning-over-MCP quality rows remain inferred from the non-MCP runs.
C. Pros & cons of each combination
The matrix above distilled into a "should I use this?" card per setup, each line grounded in the measured runs (③).
① Plain Claude
- Near-zero setup, free/cheap (≈S$0–28/mo, USD-billed)
- Best engine for analysis & drafting (top quality score, 4.70, prior-model run)
- Gets SG GST right unguarded, confirmed on the mid-2026 re-run
- You paste data by hand
- Unguarded: over-commits to customers 4/4 (promises refunds); can't file
- GST-correct today is a moving target, not a guarantee, so re-check per model
② + small-business plugin skip
- Bundles connectors (data plumbing), if you run those US apps
- All 10 finance skills carry zero SG tax logic (read directly: no GST/IRAS/CPF)
- Fabricates cash-flow confidence bands 4/4, structural to its format, not a model gap
- Its "approval gate" didn't reach the customer reply, still over-committed 4/4
- Last on quality (3.69, prior-model run); install + connector friction
③ + MCP connector live-tested
- Live auto-pull; inherits SG config (confirmed: SGD + GST 9%)
- Keeps Claude's SG knowledge, good for live-ledger reasoning
- Still can't file the F5 (no endpoint in the connector)
- No guardrail by itself, so the same behavioural risk (over-commit / invented precision)
- Plugin + MCP = fabricated bands on authoritative-looking live data
★ + SG guardrail script best Claude
- Invented forecast bands 4/4 → 0/4; over-commitment 4/4 → 0/4
- The durable win: fixes behaviour the base model doesn't, unlike tax, which it now gets on its own
- Near-zero setup (paste a recipe); works with or without plugin/MCP
- Still an LLM, so not guaranteed; can't file the F5
④ Status quo: Xero keep
- Deterministic, audit-trailed; files the F5 to IRAS (ASR+)
- Native SG GST framework (reverse charge, Customer Accounting, IGDS)
- Grant support has applied historically, so check the current scheme before budgeting
- Subscription + setup/learning curve (S$39–95/mo)
- Rigid for ad-hoc analysis & drafting (JAX emerging)
What I found: the actual run results
The most important result is a change over time. The June 2026 benchmark found unguarded Claude underpaid GST catastrophically. When I re-ran the same tests in July on a newer model, that failure had vanished, but two behavioural failures (fabricated forecast precision, over-promising to customers) persisted. That split, knowledge gaps closing while behavioural gaps stay, is the headline. Everything generated by Claude Sonnet, judged by Claude Opus (never self-graded).
the headline What changed between two model versions, and what didn't
Same tests, same synthetic data, two model versions ~two months apart. The failures that were about knowledge closed on their own; the ones about behaviour and plumbing did not.
| Failure mode | Jun 2026 prior model | Jul 2026 current model | Trajectory |
|---|---|---|---|
| Catastrophic GST underpay (importer, unguarded) | plain 40% · plugin 60% | 0 / 13 | ▼ Closing: a knowledge gap, shrinks each model release |
| Invented forecast bands (cash-flow, plugin) | 89% | 4 / 4 | ▬ Persistent: baked into the plugin's format, not knowledge |
| Over-commit to customers (complaints, unguarded) | 83% | 8 / 8 | ▬ Persistent: a behaviour, not knowledge |
| Thin SG guardrail fixes the two above | → 17% / 25% | → 0 / 4 | ✓ Durable: the guardrail earns its keep on behaviour |
| No SG connectors · can't file IRAS | gap | gap | ▬ Structural: unchanged until someone builds them |
The takeaway that won't go stale: don't buy a tool for what the base model will soon do anyway (tax knowledge). Buy guardrails for behaviour and connectors for reach, and re-run this yourself on whatever model you're on, because the top row keeps moving.
The GST safety story: a June finding that reversed in July
The realistic case is a non-expert owner who just asks "help me settle my GST." In June 2026 I ran that exact naive close 10 times per setup, on 2 GST businesses, a design studio and an importer (back-office with overseas bills), free-form, no safety hints, independently judged against the true tax figures. Unguarded Claude underpaid every time; the plugin was worst. Those June rates (prior model):
| Setup · naive prompt (10 runs per business) | Underpaid IRAS studio |
Underpaid IRAS importer |
Catastrophic¹ importer |
|---|---|---|---|
| Plain Claude no guardrails | 10 / 10 | 10 / 10 | 4 / 10 |
| + small-business plugin no guardrails | 10 / 10 | 10 / 10 | 6 / 10 |
| Plain Claude + SG script ★ | 1 / 10 | 0 / 10 | 0 / 10 |
| Plugin + SG script | 0 / 10 | 1 / 10 | 0 / 10 |
¹ Catastrophic = claimed the disallowed import GST off overseas invoices (no Customs permit) → net GST well below the true figure; the worst run underpaid by ~S$9,200. The importer is the back-office bookkeeper; the studio has no imports.
the reversal Re-run in July, the catastrophic GST error was gone
I re-ran the naive importer close on the current model. The blatant import-GST trap, claiming GST off overseas invoices without a Customs permit, fired 0 times in 13 unguarded runs (7 plain + 6 plugin). The plugin was no longer worse than plain; both produced essentially the correct net GST (~S$17,6xx). The June "guardrails are the whole game" claim for tax no longer holds: the model closed that knowledge gap on its own.
But the guardrail didn't stop earning its keep: its value moved from tax knowledge to behaviour: fabricated forecast precision and over-commitment, which did not improve (next section). Knowledge was the part the model could outgrow; behaviour isn't.
method honesty Why these are the numbers I trust
I also ran a primed version (the other 80 runs) where each run was asked "did you claim import GST off invoices?", and it scored 0% errors almost everywhere. Asking the question made the model careful and suppressed the very mistake. That 0%-vs-100% gap is itself a finding: how you prompt dominates the result. A real owner never asks themselves the safety question, so the free-form rates above are the honest ones.
Where the guardrail still earns its keep: behaviour (cash-flow & drafting)
These failures aren't about tax knowledge, they're about behaviour: inventing precision the data can't support, and promising things on the owner's behalf. Unlike GST, they did not improve on the newer model. Flaw rates, June → July re-run (July: free-form, n=4 per setup, independently judged):
| Setup | Cash-flow: invented confidence bands Jun → Jul | Complaints: unauthorised commitment¹ Jun → Jul |
|---|---|---|
| Plain Claude | 22% → 0/4 | 83% → 4/4 |
| + small-business plugin | 89% → 4/4 | 83% → 4/4 |
| + SG guardrail script | 17% → 0/4 | 25% → 0/4 |
¹ promised a refund / instalment / credit without gating to owner approval. The plugin fabricates cash-flow confidence bands in every run (4/4), a format artefact (it emits default ±30% bands), not a knowledge gap the model can outgrow. Unguarded, Claude over-commits on the owner's behalf in every complaint run (plain and plugin, 8/8), and the plugin's own "approval gate" note didn't stop it: the customer-facing reply still promised. The SG guardrail cut both to 0/4. This is the guardrail's durable value: behaviour, not tax. (Lean re-run: n=4/setup, one persona each; effects are extreme so directionally solid.)
Separately: output quality (28 expert-prompt runs, June prior-model)
With a well-written prompt that already carries the SG rules, I also scored overall quality (not just safety) across analysis, drafting and GST. Different question, different condition, and here plain Claude leads. These scores are from the June prior-model run and were not re-scored, so treat as historical; the relative ordering is the durable part.
Plain Claude won outright. The purpose-built plugin (L2) came last: its US-shaped guides added friction without quality. Adding the SG guardrail script (L3) rescued the plugin a lot, but still couldn't beat plain Claude, because the script's value isn't tied to the plugin.
Score by workflow type
| Workflow | L0 plain | L1 generic | L2 plugin | L3 plugin + SG script |
|---|---|---|---|---|
| Analysis: cash-flow snapshot | 4.72 | 4.28 | 4.02 | 4.28 |
| Drafting: customer complaints | 4.85 | 4.55 | 3.85 | 4.35 |
| Statutory: month-end + GST | 4.53 | 4.15 | 3.03 | 3.95 |
| average | 4.70 | 4.32 | 3.69 | 4.21 |
Plain Claude (L0) was best in every category. No contradiction with the safety rates above: plain Claude is the strongest engine when the prompt already carries the rules (this quality test), but on a naive prompt it has no guardrail to stop a confident underpayment. Best engine ≠ safe unguarded; that's exactly why the SG script matters.
So, do you need a "skill" or plugin?
For an expert one-shot prompt: no, plain Claude suffices.
For a non-expert SME: yes, but for behaviour, not tax. A thin SG guardrail that bans invented precision and unauthorised commitments. The current model handles SG GST itself; the guardrail's job is to stop it fabricating forecast bands and over-promising to customers (both persisted on the re-run). Not the heavy US plugin, which fabricates bands the most. Cheap, high-leverage, available today on plain Claude.
For the filed number: still buy Xero. Guardrail-Claude is safe and near-exact, but Xero gives a guaranteed figure filed straight to IRAS, with the audit trail. Claude reasons; Xero files.
evidence status What's been run
- ✓ Done: safety-rate benchmark (160 runs). The rates above come from 10 free-form runs per setup × 2 GST personas, judged by an independent model. It replaced the single-run figure with a real rate, and exposed how badly structured prompting can flatter a model (primed 0% vs free-form 100%).
- ✓ Done: SG-script config tested directly. Plain Claude + the SG script (no plugin) was measured: 0% catastrophic, matching plugin+script, confirming the script, not the plugin, is the value.
- ✓ Done: live Xero sandbox probe. Connected the Xero connector to a live Singapore org: it returned SGD + GST 9% (plus reverse-charge @9%, Customer Accounting, import/IGDS rates) purely from the org config, and USD + Canadian rates from a Global org, proving the connector inherits the account, not a country. It pulled 68 invoices live and exposed no filing endpoint. Reasoning-over-MCP quality is still inferred from the non-MCP runs.
- ✓ Done: July 2026 re-run on a newer model. Re-ran the GST-safety, cash-flow and drafting rounds. Catastrophic GST error gone (0/13 unguarded); invented bands (plugin 4/4) and over-commitment (8/8) persisted; the SG guardrail still eliminated both (0/4). The knowledge gap closed; the behavioural gaps didn't.
Still directional: 1 GST workflow, 2 personas, lean re-run (n=4–13 per cell), worth widening to more archetypes (F&B / retail / trades) and a field test on a real consented SME's live Xero. And re-running per model release, since the knowledge-gap results keep moving.
Coverage: what was tested vs reasoned
Honesty check: I did not test every cell. This is the per-setup view of the four rounds in "What was tested" (tab ①): ✓ where measured, ◐ where reasoned / web-verified, ✗ where untested.
| Setup | Analysis | Drafting | Statutory | Safety rate free-form | Live data (MCP) |
|---|---|---|---|---|---|
| Plain Claude | ✓ | ✓ | ✓ | ✓ | ✗ |
| Generic skills | ✓ | ✓ | ✓ | – | ✗ |
| + plugin | ✓ | ✓ | ✓ | ✓ | ✗ |
| + plugin + SG script | ✓ | ✓ | ✓ | ✓ | ✗ |
| Plain + SG script | ◐ safety | ◐ safety | – | ✓ | ✗ |
| + MCP (no plugin) | ✗ | ✗ | ✗ | ✗ | ✓ live |
| + plugin + MCP | ✗ | ✗ | ✗ | ✗ | ◐ reasoned |
| Status quo (Xero) | n/a | n/a | ◐ web | – | ◐ web |
✓ measured · – not run (*inferable from plugin+script, which it matched on statutory) · ◐ reasoned / web-verified · ✗ untested. The live-data (MCP) connector layer is now measured (SG org → SGD + GST 9%, live pull, no filing), and free-form safety rates now span cash-flow, drafting & statutory. Remaining gap: reasoning-over-MCP quality (analysis/drafting via live data).
Method & caveats: Directional, not exhaustive: 5 synthetic (PDPA-safe) personas, 3 of 8 planned workflows. Runs: 28 expert quality (round 1) + 160 GST safety-benchmark (round 2, free-form & primed) + 90 cash-flow/drafting safety (round 3) + 6 preliminary naive ≈ 280 (Jun). Plus a July 2026 re-run on a newer Sonnet (~37 runs: importer GST ×13, cash-flow ×12, drafting ×12); the Jun→Jul comparison across model versions is the point, not a like-for-like control. Generated by Claude Sonnet, judged by Claude Opus (never the same model grading its own output). The "+ MCP" cells are reasoned from the measured layers + web-verified connector capabilities, not separately live-run. Connector facts verified May–Jun 2026. Plugin scope: the small-business plugin has 31 skills. I ran 3 (one per archetype) and read all 10 finance skills (0 contain GST/IRAS/CPF); its ~18 CRM/marketing/ops/contract/hiring skills were not evaluated.
Tool-by-tool: which SG SME surfaces an AI can actually reach
"Can Claude help with X?" depends on whether a connector (MCP) exists for the tool the SME uses, what that connector is allowed to do, and the CorpPass gate on government filing. Here's every surface a typical SG SME touches, the detail behind the Data layer in ② and the "+ MCP" column.
| Surface (what the SME does) | What SG SMEs use | Connector today? | What AI can do | Workaround now |
|---|---|---|---|---|
| Accounting & books | Xero (most common), QuickBooks, MYOB | ✓ Xero official · QBO sandbox 📦 QBO | Read + post invoices, bills, ledger | Wire the Xero connector; reason in Claude |
| GST F5 filing | myTax Portal + CorpPass | ✗ no IRAS connector | Prepare figures only, can't file | Xero one-click (ASR+); or key into myTax; or tax agent |
| Getting paid | PayNow · NETS · FAST (+ cards) | ✗ rails none · Stripe/PayPal ✓ 📦 · Square N/A in SG | Card/online payments & refunds; not local rails | Local gateway (e.g. HitPay); reconcile in Xero |
| Customer comms | WhatsApp Business (top channel, industry est.) | ~ no official Meta one · third-party wrappers 📦 Slack/Gmail | Read chats, draft replies | Third-party wrapper, or paste thread → Claude |
| Selling: marketplaces | Shopee · Lazada · Carousell | ✗ no seller connector (scrapers only) | Read public listings only | Export seller-centre CSV → Claude |
| Delivery: F&B | GrabFood · Foodpanda | ✗ no merchant connector | – | Export from merchant portal → Claude |
| Point of sale | StoreHub · Eats365 | ~ StoreHub community · Eats365 none | Read sales / inventory | Community connector or API export |
| Payroll / CPF | Talenox, local HRMS | ✗ none | Payslip/CPF calc checks only | Software files CPF; Claude sense-checks |
| Entity admin | ACRA / BizFile (CorpPass) | ~ read registry (open data) · no write | Look up companies / UEN / GST status | File in BizFile; Claude preps |
| Office / marketing | Canva · Google / M365 · Gmail | ✓ official 📦 in plugin | Full: read + create | Use as-is: the clean fits |
the real gap The surfaces that define an SG SME's day have no connector
Getting paid (PayNow/NETS), talking to customers (WhatsApp Business), selling (Shopee/Grab) and filing (IRAS), the four things an SG SME does most, have no usable AI connector today, from the plugin or anywhere. The plugin bundles the office tools and US payment apps, which are the least SG-critical. An SG-shaped connector layer (IRAS, PayNow, WhatsApp Business, local POS/marketplaces) is the highest-leverage thing a programme could build.