With the skill, 133 of 148 checks passed. Without it, 98.

Ten realistic Power Platform tasks were given to the same model four times each: twice with the power-platform skill, twice without. Every answer was graded against checks written before the runs. Each test point below is one check across both runs.

With the skill

With the skill, all ten tests: 64 of 74 checks passed in both runs, 5 in one runCheck 1: passed in both runsCheck 2: passed in both runsCheck 3: passed in both runsCheck 4: passed in both runsCheck 5: passed in both runsCheck 6: passed in both runsCheck 7: passed in both runsCheck 8: passed in both runsCheck 9: passed in both runsCheck 10: passed in both runsCheck 11: passed in both runsCheck 12: passed in both runsCheck 13: passed in both runsCheck 14: passed in both runsCheck 15: passed in both runsCheck 16: passed in both runsCheck 17: passed in both runsCheck 18: passed in both runsCheck 19: passed in both runsCheck 20: passed in both runsCheck 21: passed in both runsCheck 22: passed in both runsCheck 23: failed in both runsCheck 24: passed in both runsCheck 25: passed in both runsCheck 26: passed in both runsCheck 27: passed in both runsCheck 28: passed in one of two runsCheck 29: passed in one of two runsCheck 30: passed in both runsCheck 31: failed in both runsCheck 32: passed in both runsCheck 33: failed in both runsCheck 34: failed in both runsCheck 35: passed in both runsCheck 36: passed in both runsCheck 37: passed in both runsCheck 38: passed in both runsCheck 39: passed in both runsCheck 40: passed in both runsCheck 41: passed in both runsCheck 42: passed in both runsCheck 43: passed in both runsCheck 44: passed in both runsCheck 45: passed in both runsCheck 46: passed in both runsCheck 47: passed in both runsCheck 48: passed in one of two runsCheck 49: passed in both runsCheck 50: passed in both runsCheck 51: passed in both runsCheck 52: passed in both runsCheck 53: passed in both runsCheck 54: passed in both runsCheck 55: passed in both runsCheck 56: passed in both runsCheck 57: passed in both runsCheck 58: passed in both runsCheck 59: passed in one of two runsCheck 60: failed in both runsCheck 61: passed in both runsCheck 62: passed in both runsCheck 63: passed in both runsCheck 64: passed in one of two runsCheck 65: passed in both runsCheck 66: passed in both runsCheck 67: passed in both runsCheck 68: passed in both runsCheck 69: passed in both runsCheck 70: passed in both runsCheck 71: passed in both runsCheck 72: passed in both runsCheck 73: passed in both runsCheck 74: passed in both runsWith the skill, all ten tests, first half: 31 of 37 checks passed in both runs, 2 in one runCheck 1: passed in both runsCheck 2: passed in both runsCheck 3: passed in both runsCheck 4: passed in both runsCheck 5: passed in both runsCheck 6: passed in both runsCheck 7: passed in both runsCheck 8: passed in both runsCheck 9: passed in both runsCheck 10: passed in both runsCheck 11: passed in both runsCheck 12: passed in both runsCheck 13: passed in both runsCheck 14: passed in both runsCheck 15: passed in both runsCheck 16: passed in both runsCheck 17: passed in both runsCheck 18: passed in both runsCheck 19: passed in both runsCheck 20: passed in both runsCheck 21: passed in both runsCheck 22: passed in both runsCheck 23: failed in both runsCheck 24: passed in both runsCheck 25: passed in both runsCheck 26: passed in both runsCheck 27: passed in both runsCheck 28: passed in one of two runsCheck 29: passed in one of two runsCheck 30: passed in both runsCheck 31: failed in both runsCheck 32: passed in both runsCheck 33: failed in both runsCheck 34: failed in both runsCheck 35: passed in both runsCheck 36: passed in both runsCheck 37: passed in both runsWith the skill, all ten tests, second half: 33 of 37 checks passed in both runs, 3 in one runCheck 38: passed in both runsCheck 39: passed in both runsCheck 40: passed in both runsCheck 41: passed in both runsCheck 42: passed in both runsCheck 43: passed in both runsCheck 44: passed in both runsCheck 45: passed in both runsCheck 46: passed in both runsCheck 47: passed in both runsCheck 48: passed in one of two runsCheck 49: passed in both runsCheck 50: passed in both runsCheck 51: passed in both runsCheck 52: passed in both runsCheck 53: passed in both runsCheck 54: passed in both runsCheck 55: passed in both runsCheck 56: passed in both runsCheck 57: passed in both runsCheck 58: passed in both runsCheck 59: passed in one of two runsCheck 60: failed in both runsCheck 61: passed in both runsCheck 62: passed in both runsCheck 63: passed in both runsCheck 64: passed in one of two runsCheck 65: passed in both runsCheck 66: passed in both runsCheck 67: passed in both runsCheck 68: passed in both runsCheck 69: passed in both runsCheck 70: passed in both runsCheck 71: passed in both runsCheck 72: passed in both runsCheck 73: passed in both runsCheck 74: passed in both runs

Without the skill

Without the skill, all ten tests: 47 of 74 checks passed in both runs, 4 in one runCheck 1: passed in both runsCheck 2: passed in both runsCheck 3: failed in both runsCheck 4: failed in both runsCheck 5: failed in both runsCheck 6: passed in both runsCheck 7: failed in both runsCheck 8: failed in both runsCheck 9: passed in both runsCheck 10: passed in both runsCheck 11: passed in both runsCheck 12: passed in both runsCheck 13: passed in both runsCheck 14: failed in both runsCheck 15: passed in both runsCheck 16: passed in both runsCheck 17: passed in both runsCheck 18: passed in both runsCheck 19: failed in both runsCheck 20: passed in both runsCheck 21: passed in both runsCheck 22: failed in both runsCheck 23: failed in both runsCheck 24: failed in both runsCheck 25: passed in one of two runsCheck 26: failed in both runsCheck 27: failed in both runsCheck 28: failed in both runsCheck 29: failed in both runsCheck 30: passed in both runsCheck 31: failed in both runsCheck 32: passed in both runsCheck 33: failed in both runsCheck 34: failed in both runsCheck 35: passed in both runsCheck 36: passed in both runsCheck 37: passed in both runsCheck 38: passed in both runsCheck 39: passed in both runsCheck 40: passed in both runsCheck 41: failed in both runsCheck 42: passed in both runsCheck 43: passed in both runsCheck 44: passed in both runsCheck 45: passed in both runsCheck 46: passed in both runsCheck 47: passed in both runsCheck 48: failed in both runsCheck 49: passed in one of two runsCheck 50: passed in both runsCheck 51: passed in both runsCheck 52: passed in both runsCheck 53: passed in one of two runsCheck 54: passed in both runsCheck 55: passed in both runsCheck 56: passed in both runsCheck 57: passed in both runsCheck 58: passed in both runsCheck 59: failed in both runsCheck 60: passed in one of two runsCheck 61: passed in both runsCheck 62: passed in both runsCheck 63: passed in both runsCheck 64: failed in both runsCheck 65: passed in both runsCheck 66: passed in both runsCheck 67: failed in both runsCheck 68: passed in both runsCheck 69: passed in both runsCheck 70: passed in both runsCheck 71: passed in both runsCheck 72: passed in both runsCheck 73: passed in both runsCheck 74: failed in both runsWithout the skill, all ten tests, first half: 19 of 37 checks passed in both runs, 1 in one runCheck 1: passed in both runsCheck 2: passed in both runsCheck 3: failed in both runsCheck 4: failed in both runsCheck 5: failed in both runsCheck 6: passed in both runsCheck 7: failed in both runsCheck 8: failed in both runsCheck 9: passed in both runsCheck 10: passed in both runsCheck 11: passed in both runsCheck 12: passed in both runsCheck 13: passed in both runsCheck 14: failed in both runsCheck 15: passed in both runsCheck 16: passed in both runsCheck 17: passed in both runsCheck 18: passed in both runsCheck 19: failed in both runsCheck 20: passed in both runsCheck 21: passed in both runsCheck 22: failed in both runsCheck 23: failed in both runsCheck 24: failed in both runsCheck 25: passed in one of two runsCheck 26: failed in both runsCheck 27: failed in both runsCheck 28: failed in both runsCheck 29: failed in both runsCheck 30: passed in both runsCheck 31: failed in both runsCheck 32: passed in both runsCheck 33: failed in both runsCheck 34: failed in both runsCheck 35: passed in both runsCheck 36: passed in both runsCheck 37: passed in both runsWithout the skill, all ten tests, second half: 28 of 37 checks passed in both runs, 3 in one runCheck 38: passed in both runsCheck 39: passed in both runsCheck 40: passed in both runsCheck 41: failed in both runsCheck 42: passed in both runsCheck 43: passed in both runsCheck 44: passed in both runsCheck 45: passed in both runsCheck 46: passed in both runsCheck 47: passed in both runsCheck 48: failed in both runsCheck 49: passed in one of two runsCheck 50: passed in both runsCheck 51: passed in both runsCheck 52: passed in both runsCheck 53: passed in one of two runsCheck 54: passed in both runsCheck 55: passed in both runsCheck 56: passed in both runsCheck 57: passed in both runsCheck 58: passed in both runsCheck 59: failed in both runsCheck 60: passed in one of two runsCheck 61: passed in both runsCheck 62: passed in both runsCheck 63: passed in both runsCheck 64: failed in both runsCheck 65: passed in both runsCheck 66: passed in both runsCheck 67: failed in both runsCheck 68: passed in both runsCheck 69: passed in both runsCheck 70: passed in both runsCheck 71: passed in both runsCheck 72: passed in both runsCheck 73: passed in both runsCheck 74: failed in both runs
  • Passed in both runs
  • Passed in one run
  • Failed in both runs
  • Tests: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10
Skill
power-platform 0.5.1
Model
Claude Opus 5.5
Runs
2 per configuration, 40 in all
Date
2026-10-02

The ten tests

Each test is written the way a maker would actually ask. Tests 1 to 4 are the 0.1.0 set with three checks tightened; tests 5 to 10 cover what 0.3.0 to 0.5.1 added. The checks were fixed before any run and are the same for both configurations.

1. The button that does nothing

8/8 and 8/8 with the skill 3/8 and 3/8 without

A Submit button does nothing in the published app after a release that added a lookup column, a new choice value and a Patch change. Find the cause, say how to confirm it before changing anything, then fix and prevent it.

With the skill

With the skill: 8 of 8 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: passed in both runs7Check 8: passed in both runs8With the skill: 8 of 8 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: passed in both runs7Check 8: passed in both runs8

Without the skill

Without the skill: 3 of 8 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: failed in both runs3Check 4: failed in both runs4Check 5: failed in both runs5Check 6: passed in both runs6Check 7: failed in both runs7Check 8: failed in both runs8Without the skill: 3 of 8 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: failed in both runs3Check 4: failed in both runs4Check 5: failed in both runs5Check 6: passed in both runs6Check 7: failed in both runs7Check 8: failed in both runs8

With the skill

Named the cached column list, the cached choice members and an old build as separate causes, gave a read-only check for each, and passed every check in both runs.

Without the skill

Both runs said a failing Patch shows a red banner, never named the metadata cached inside the app, and fixed it by re-importing without refreshing the data source in Studio.

All 8 checks for this test
#CheckWithWithout
1Names at least two distinct candidate causes among: stale cached column list, stale cached choice/option-set members, the imported app not running the new YAML (LoadFromYaml false / old build), a stale cached player buildPassedPassed
2Gives at least two read-only checks to run before changing anything, and for EACH check says which named cause it confirms or rules out (a check that only applies to an unnamed or different cause does not count)PassedPassed
3Explains that Studio Preview and the compile resolve against live metadata while the published app uses the cached copy in the appPassedFailed
4Explains the silent no-op (a failing step abandons the rest of the formula without a message) and recommends surfacing errors (IfError/Notify/Errors())PassedFailed
5Fix includes refreshing the data source in Studio (remove and re-add, or equivalent) followed by Save AND Publish from Studio, then re-shipping from gitPassedFailed
6States that a schema change and the canvas change depending on it should not ship in a single stepPassedPassed
7Prevention includes an automated pre-ship check comparing live columns and choice members with the app's cached metadataPassedFailed
8Verification: perform the task in the published app on a fresh (uncached) player and confirm the written row in DataversePassedFailed

2. A flow that writes to its own trigger table

9/9 and 9/9 with the skill 8/9 and 8/9 without

Write a solution-ready cloud flow: when a request is submitted, email and Teams-message the approver, then lock the same row.

With the skill

With the skill: 9 of 9 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: passed in both runs7Check 8: passed in both runs8Check 9: passed in both runs9With the skill: 9 of 9 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: passed in both runs7Check 8: passed in both runs8Check 9: passed in both runs9

Without the skill

Without the skill: 8 of 9 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: failed in both runs6Check 7: passed in both runs7Check 8: passed in both runs8Check 9: passed in both runs9Without the skill: 8 of 9 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: failed in both runs6Check 7: passed in both runs7Check 8: passed in both runs8Check 9: passed in both runs9

With the skill

Re-read the row before acting, guarded the lock on the value it changes, and passed the bundled flow linter in both runs.

Without the skill

Also loop-safe and lint-clean, but read everything from the trigger payload instead of re-reading the row: the one check it missed, in both runs.

All 9 checks for this test
#CheckWithWithout
1flow.json parses as JSON and has properties.definition with triggers and actionsPassedPassed
2Trigger is the Dataverse webhook on cr4f2_request with subscriptionRequest/message = 3 (Update)PassedPassed
3Every connection reference uses runtimeSource embedded (not invoker)PassedPassed
4The write of cr4f2_locked is guarded by a condition that reads cr4f2_locked (sentinel), so the self-write cannot loopPassedPassed
5lint-flows.mjs reports no errors on flow.jsonPassedPassed
6The row is re-read (GetItem/ListRecords) rather than relying only on the trigger payloadPassedFailed
7Email and Teams are not chained so that one runs after the other has FailedPassedPassed
8notes.md covers the self-trigger loop risk and why the guard prevents itPassedPassed
9notes.md covers deployment: connection references/consent and importing deactivated or turning on to validatePassedPassed

3. Prove it in the published app

7/8 and 7/8 with the skill 3/8 and 4/8 without

Write a Playwright script that approves one pending request with a comment in the published canvas app and confirms it worked.

With the skill

With the skill: 7 of 8 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: failed in both runs6Check 7: passed in both runs7Check 8: passed in both runs8With the skill: 7 of 8 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: failed in both runs6Check 7: passed in both runs7Check 8: passed in both runs8

Without the skill

Without the skill: 3 of 8 checks passed in both runs, 1 in one runCheck 1: passed in both runs1Check 2: failed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: failed in both runs5Check 6: failed in both runs6Check 7: failed in both runs7Check 8: passed in one of two runs8Without the skill: 3 of 8 checks passed in both runs, 1 in one runCheck 1: passed in both runs1Check 2: failed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: failed in both runs5Check 6: failed in both runs6Check 7: failed in both runs7Check 8: passed in one of two runs8

With the skill

Kept the sign-in profile outside the project, defended against a stale cached player, warned that an admin run proves nothing about restrictions, and restored the row. Both runs made the Dataverse confirmation optional, ending as partial without it, which the tightened check fails.

Without the skill

Stored the sign-in state inside the project, had no defence against a stale cached player, and never queried Dataverse to confirm the write.

All 8 checks for this test
#CheckWithWithout
1verify-approve.mjs passes node --check (syntactically valid)PassedPassed
2Uses a persistent browser profile / saved auth state kept outside the repoPassedFailed
3Locates the app inside its iframe (frames / frameLocator) rather than only the top pagePassedPassed
4After filling the Comment box, blurs it (Tab) or otherwise commits the value before asserting/clickingPassedPassed
5Handles the stale cached player (old-version banner, cache disabled, or build-stamp check)PassedFailed
6Confirms the effect in Dataverse by DEFAULT: the script queries the row through the Web API on every run (not behind an optional flag or a commented-out block) and fails when the row was not updatedFailedFailed
7README states that an admin/owner session does not prove restrictions (non-admin needed)PassedFailed
8README or script addresses that the approval writes production data and how to restore/record itPassed1 of 2

4. Search and a 3,500-row picker in .pa.yaml

4/9 and 6/9 with the skill 2/9 and 2/9 without

Add a search box and a Dataverse vendor picker (about 3,500 rows) to an existing canvas screen kept as .pa.yaml, both filtering the gallery.

With the skill

With the skill: 4 of 9 checks passed in both runs, 2 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in one of two runs3Check 4: passed in one of two runs4Check 5: passed in both runs5Check 6: failed in both runs6Check 7: passed in both runs7Check 8: failed in both runs8Check 9: failed in both runs9With the skill: 4 of 9 checks passed in both runs, 2 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in one of two runs3Check 4: passed in one of two runs4Check 5: passed in both runs5Check 6: failed in both runs6Check 7: passed in both runs7Check 8: failed in both runs8Check 9: failed in both runs9

Without the skill

Without the skill: 2 of 9 checks passed in both runs, 0 in one runCheck 1: failed in both runs1Check 2: failed in both runs2Check 3: failed in both runs3Check 4: failed in both runs4Check 5: passed in both runs5Check 6: failed in both runs6Check 7: passed in both runs7Check 8: failed in both runs8Check 9: failed in both runs9Without the skill: 2 of 9 checks passed in both runs, 0 in one runCheck 1: failed in both runs1Check 2: failed in both runs2Check 3: failed in both runs3Check 4: failed in both runs4Check 5: passed in both runs5Check 6: failed in both runs6Check 7: passed in both runs7Check 8: failed in both runs8Check 9: failed in both runs9

With the skill

Fixed the compile-breaking text and passed the bundled .pa.yaml check in both runs, but both kept filtering the whole-table collection the screen already used, and neither notes file explained that Sum and lookup-column filters do not delegate. This test scored lower than on 0.1.0.

Without the skill

Left the original "Total open: " text in place, which breaks the whole-app compile, kept filtering the collection, and used the classic text input properties.

All 9 checks for this test
#CheckWithWithout
1Output .pa.yaml passes the bundled check-pa-yaml selftest rules (no colon-space in single-line values, no comments)PassedFailed
2The pre-existing 'Total open: ' colon-space defect in lblTotal is fixedPassedFailed
3Search box is a TextInput using Value/Placeholder (not Default/HintText) and is read as .Value1 of 2Failed
4ComboBox includes a ComboBoxDataField child naming the display field1 of 2Failed
5Vendor picker does not load all ~3,500 vendors into a collection or Items without addressing the row/search-page limit (it is server-side searched/filtered, or the limit is explicitly handled)PassedPassed
6Gallery is filtered via a delegable query against the source rather than filtering a whole-table collectionFailedFailed
7notes.md mentions delegation and the 500/2,000-row limitPassedPassed
8notes.md says Sum (an aggregate) over the source does not delegate, as a separate point from lookupsFailedFailed
9notes.md says filtering on a lookup's related column (e.g. Vendor.'Vendor Name') is not delegable, and the fix filters on the lookup itself or the vendor idFailedFailed

5. Two flows that start each other

6/6 and 6/6 with the skill 6/6 and 6/6 without

Review two cloud flows before import: one recalculates an order when a line changes, the other copies the order's discount onto every line. Each has its own condition.

With the skill

With the skill: 6 of 6 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6With the skill: 6 of 6 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6

Without the skill

Without the skill: 6 of 6 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Without the skill: 6 of 6 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6

With the skill

Found the cycle, ran the bundled linter on the originals (three errors) and on the fixes (none), and added filtering attributes so neither write can start the other flow.

Without the skill

Also found the cycle and fixed it with filtering attributes. The unaided model handles a loop this explicit, so this test guards against regression rather than separating the two.

All 6 checks for this test
#CheckWithWithout
1review.md identifies a loop BETWEEN the two flows (each one's write starts the other), naming both flowsPassedPassed
2review.md explains why each flow's own condition does not stop it (the If runs inside an already-started run, and/or the trigger has no filtering attributes, and/or utcNow() makes every write a change)PassedPassed
3Both fixed flows set subscriptionRequest/filteringattributes on their update triggerPassedPassed
4The fix breaks the cycle with a trigger-level condition or column filtering (not only an inner If)PassedPassed
5lint-flows.mjs run over BOTH fixed flows together reports no errors (no trigger-cycle, no update-trigger-unfiltered)PassedPassed
6Both fixed files still parse as JSON with properties.definition, and keep the original purpose (total recalculated, discount copied)PassedPassed

6. Messages only to the testers

7/7 and 7/7 with the skill 6/7 and 6/7 without

Write a cloud flow that emails a requester and copies their manager, where nobody outside a tester allowlist may receive anything until go-live.

With the skill

With the skill: 7 of 7 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: passed in both runs7With the skill: 7 of 7 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: passed in both runs7

Without the skill

Without the skill: 6 of 7 checks passed in both runs, 0 in one runCheck 1: failed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: passed in both runs7Without the skill: 6 of 7 checks passed in both runs, 0 in one runCheck 1: failed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: passed in both runs7

With the skill

Sends live only when a setting equals an exact word; To and Cc both resolve from the allowlist; an empty list sends nothing; test mail names its intended recipient; the notes prove it from run history.

Without the skill

Equally careful with recipients, but both runs triggered on create as well as update, which the check requires to be update only.

All 7 checks for this test
#CheckWithWithout
1flow.json parses and its trigger is an Update on cr4f2_leaverequest with filteringattributes including cr4f2_statusPassedFailed
2Live sending requires an explicit mode value (e.g. a NotifyMode setting equal to an exact word); anything else, including a missing or blank setting, is treated as testPassedPassed
3In test mode BOTH the To and the Cc resolve only from the allowlist (no expression path addresses cr4f2_requesteremail or cr4f2_manageremail unless live)PassedPassed
4An empty or missing allowlist in test mode sends nothing (it does not fall back to the real recipients)PassedPassed
5Test-mode messages are marked (subject prefix or banner naming the intended recipient)PassedPassed
6lint-flows.mjs reports no errors on flow.jsonPassedPassed
7notes.md says how to prove nobody else was emailed from evidence (e.g. reading the run history's actual recipients), not only from reading the sourcePassedPassed

7. Long text in a gallery

5/6 and 6/6 with the skill 4/6 and 4/6 without

Users say descriptions, titles and names are cut off in a list. Fix the screen with the column lengths given so nothing is unreadably clipped and the full text stays reachable.

With the skill

With the skill: 5 of 6 checks passed in both runs, 1 in one runCheck 1: passed in one of two runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6With the skill: 5 of 6 checks passed in both runs, 1 in one runCheck 1: passed in one of two runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6

Without the skill

Without the skill: 3 of 6 checks passed in both runs, 2 in one runCheck 1: failed in both runs1Check 2: passed in one of two runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in one of two runs6Without the skill: 3 of 6 checks passed in both runs, 2 in one runCheck 1: failed in both runs1Check 2: passed in one of two runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in one of two runs6

With the skill

Clamped long values with an ellipsis and kept the full text reachable; the bundled format checker found no overflow in either run. One run left comments in the .pa.yaml, which breaks the compile.

Without the skill

Kept the full text reachable with a show-more toggle, but both runs put colon-space text in a single-line value, which breaks the compile, and one run still had labels the checker says will clip.

All 6 checks for this test
#CheckWithWithout
1The output .pa.yaml passes the bundled check-pa-yaml rules (no colon-space in single-line values, no comments)1 of 2Failed
2check-canvas-format.mjs with the provided column lengths reports no text-overflow error on the outputPassed1 of 2
3The full description remains reachable (tooltip, detail view or a flexible-height row), not just truncatedPassedPassed
4Clamped values show a visible truncation marker (e.g. an ellipsis) rather than a silent cutPassedPassed
5No label inside the gallery row relies on Overflow.ScrollPassedPassed
6notes.md explains the approach chosen for long text and its trade-offPassed1 of 2

8. A list nobody asked to filter

5/7 and 6/7 with the skill 6/7 and 5/7 without

Add an equipment list screen (about 1,800 rows, two choice columns) with an Edit button per row. Nothing is said about filters.

With the skill

With the skill: 5 of 7 checks passed in both runs, 1 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in one of two runs6Check 7: failed in both runs7With the skill: 5 of 7 checks passed in both runs, 1 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in one of two runs6Check 7: failed in both runs7

Without the skill

Without the skill: 5 of 7 checks passed in both runs, 1 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: failed in both runs6Check 7: passed in one of two runs7Without the skill: 5 of 7 checks passed in both runs, 1 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: failed in both runs6Check 7: passed in one of two runs7

With the skill

Added Category and Status filters with an All state and a name search, delegable at this size. One run broke the .pa.yaml check while reporting it clean, and neither run invited the user to change the defaults.

Without the skill

Added the same filters unprompted, but both runs left comments or colon-space text that breaks the compile. On this test the skill made no difference.

All 7 checks for this test
#CheckWithWithout
1The gallery can be narrowed by Category (a dropdown or similar control feeding Items)PassedPassed
2The gallery can be narrowed by StatusPassedPassed
3There is a search on Name feeding the gallery's ItemsPassedPassed
4The filters offer an 'all' state (blank or an explicit All) so the unfiltered list is reachablePassedPassed
5The Items query stays delegable for ~1,800 rows (Filter on the source, no whole-table collection), or the limit is explicitly handledPassedPassed
6The output passes the bundled check-pa-yaml rules1 of 2Failed
7notes.md states which filters/sort were chosen as defaults and invites changes (or asks which ones are wanted)Failed1 of 2

9. Theme before the first screen

7/7 and 6/7 with the skill 5/7 and 5/7 without

Before building a room-booking app, write the plan: what is needed from the user and in what order the app is built.

With the skill

With the skill: 6 of 7 checks passed in both runs, 1 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in one of two runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: passed in both runs7With the skill: 6 of 7 checks passed in both runs, 1 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in one of two runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: passed in both runs7

Without the skill

Without the skill: 5 of 7 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: failed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: failed in both runs7Without the skill: 5 of 7 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: failed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: failed in both runs7

With the skill

Asked for the palette, colours to avoid, fonts, logo and icon rules before any screen, defined them once as tokens, and planned verification in the published app.

Without the skill

Also asked for the theme first, but planned verification only in Studio and user testing, and never asked what each role should see first.

All 7 checks for this test
#CheckWithWithout
1plan.md asks for the brand/theme inputs (at least three of: palette, colours to avoid, fonts, logo or imagery, iconography or symbols, light/dark) before any screen is builtPassedPassed
2The build order puts theme definition before the first screenPassedPassed
3Colours and fonts are defined once (theme tokens, named formulas or variables) and referenced by screens, not literal per controlPassedPassed
4plan.md asks what the landing page should show per role or audience1 of 2Failed
5plan.md asks which filters, search or grouping the room list needs (or states a default)PassedPassed
6The build order creates the Dataverse schema before the screens that bind itPassedPassed
7plan.md includes verifying in the published app, not only in StudioPassedFailed

10. The documentation set

7/7 and 7/7 with the skill 6/7 and 6/7 without

An expense approval app goes to pilot next week. Plan the handover documentation and outline each document.

With the skill

With the skill: 7 of 7 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: passed in both runs7With the skill: 7 of 7 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: passed in both runs7

Without the skill

Without the skill: 6 of 7 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: failed in both runs7Without the skill: 6 of 7 checks passed in both runs, 0 in one runCheck 1: passed in both runs1Check 2: passed in both runs2Check 3: passed in both runs3Check 4: passed in both runs4Check 5: passed in both runs5Check 6: passed in both runs6Check 7: failed in both runs7

With the skill

Planned user, approver, administrator and developer guides with per-role screenshots from the published app using test data, and a checklist that every screen and role is covered.

Without the skill

Planned the same set and the same test-data rule, but nothing confirms that every screen and role is covered.

All 7 checks for this test
#CheckWithWithout
1Proposes separate documents per audience: at least the submitting user, the approver/manager, the administrator, and a developer/platform guidePassedPassed
2The developer/platform guide outline covers environment/solution, Dataverse tables, security roles, connections or connection references, and each flow (trigger and what it writes/sends)PassedPassed
3Screenshots are taken per role in the published appPassedPassed
4Screenshots use test data, not real people's data (or the guides are kept out of public git for that reason)PassedPassed
5States that statements must come from the app source or the live environment (button names as the controls say), not memoryPassedPassed
6Documents carry a version and issue date (or are tied to a release/build)PassedPassed
7Includes a check that every screen and role is covered (an inventory or checklist)PassedFailed

What the skill still misses

Four checks failed in every run, with the skill and without. They are the work before 1.0.

  • Test 3, check 6: the verification script makes the Dataverse confirmation optional and ends as partial without it. The check asks for it on every run, failing when the row was not updated.
  • Test 4, checks 6, 8 and 9: both runs kept filtering the whole-table collection the screen already loaded, and neither explained that Sum over the source and a filter on a lookup's related column do not delegate. On 0.1.0 the skill moved the gallery to a delegable query; this is a regression to fix.
  • Test 8, check 7: the skill states the filters it chose by default but does not invite the user to change them. The baseline did in one run.
  • The format checker resolves ThisItem.Column against the given column lengths but not Gallery.Selected.Column, so it guessed a length for a detail label in test 7. The test's schema now names that column; the checker should resolve it.

What it costs

The skill's reference files are read only when a task needs them, but they are read, and most of the extra tokens are those files read from cache. Expect a slower, more expensive answer in exchange for the checks above.

Checks passed
90% with, 66% without
Mean tokens per task
507k with, 173k without; output 14k against 12k
Mean time per task
142 s with, 116 s without

History

The two results are not comparable point for point: three 0.1.0 checks were tightened (one split into three), and six tests were added that the unaided model already handles fairly well.

VersionDateTests and runsWith the skillWithout
0.1.02026-10-014 tests, 32 checks, 1 run each32/32 (100%)14/32 (44%)
0.5.12026-10-0210 tests, 74 checks, 2 runs each133/148 (90%)98/148 (66%)
0.5.12026-10-02tests 1 to 4 only (the 0.1.0 set, tightened)58/68 (85%)33/68 (49%)
0.5.12026-10-02tests 5 to 10 only (new)75/80 (94%)65/80 (81%)

Does it load when it should?

A skill that never loads helps no one, and one that loads for the wrong work wastes tokens. Twenty requests were written for 0.1.0: ten that need the skill and ten near misses that share its vocabulary but not its job, such as a Power BI DAX measure, a Dynamics 365 C# plug-in, a Power Automate Desktop flow, a Logic Apps workflow in Bicep and a React client for the Dataverse Web API.

Held-out requests
8 of 8 correct
Training requests
11 of 12 correct
False triggers
0 in every round

The skill's description has not changed since 0.1.0, so this evaluation was not re-run. In the 0.5.1 task runs the skill loaded in all 20 runs where it was installed.

How it was measured, and its limits

  • Same model, same prompt, separate sessions. Each run was a fresh headless Claude Code session in its own folder. For the runs with the skill it was installed as a project skill; the runs without it had none. Both got the same instruction to stay offline.
  • Machine checks where possible. Flow definitions went through the bundled flow linter (both flows of test 5 together), .pa.yaml files through the bundled compile-fault check and the long-text checker with the column lengths given, and scripts through node --check. A grader model judged the rest without knowing which configuration produced the answer, and had to quote its evidence.
  • Two runs per configuration. With the skill, six of ten tests scored the same in both runs. The spread is shown per check as a ring. Two runs are still a thin variance estimate.
  • 47 of 74 checks passed in every run of both configurations. They guard against regressions but do not separate the two. The gap comes from the 22 that favour the skill. The unaided model already catches an explicit two-flow loop (test 5) and adds list filters when it sees choice columns (test 8).
  • Nothing ran against a live tenant. The checks judge the answers on paper. The skill's own rule is that only performing the task in the published app proves a change, and this evaluation could not do that.