The ten tests
Each test is written the way a maker would actually ask. Tests 1 to 4 are the 0.1.0 set with three checks tightened; tests 5 to 10 cover what 0.3.0 to 0.5.1 added. The checks were fixed before any run and are the same for both configurations.
1. The button that does nothing
8/8 and 8/8 with the skill 3/8 and 3/8 without
A Submit button does nothing in the published app after a release that added a lookup column, a new choice value and a Patch change. Find the cause, say how to confirm it before changing anything, then fix and prevent it.
With the skill
Without the skill
With the skill
Named the cached column list, the cached choice members and an old build as separate causes, gave a read-only check for each, and passed every check in both runs.
Without the skill
Both runs said a failing Patch shows a red banner, never named the metadata cached inside the app, and fixed it by re-importing without refreshing the data source in Studio.
All 8 checks for this test
| # | Check | With | Without |
|---|---|---|---|
| 1 | Names at least two distinct candidate causes among: stale cached column list, stale cached choice/option-set members, the imported app not running the new YAML (LoadFromYaml false / old build), a stale cached player build | Passed | Passed |
| 2 | Gives at least two read-only checks to run before changing anything, and for EACH check says which named cause it confirms or rules out (a check that only applies to an unnamed or different cause does not count) | Passed | Passed |
| 3 | Explains that Studio Preview and the compile resolve against live metadata while the published app uses the cached copy in the app | Passed | Failed |
| 4 | Explains the silent no-op (a failing step abandons the rest of the formula without a message) and recommends surfacing errors (IfError/Notify/Errors()) | Passed | Failed |
| 5 | Fix includes refreshing the data source in Studio (remove and re-add, or equivalent) followed by Save AND Publish from Studio, then re-shipping from git | Passed | Failed |
| 6 | States that a schema change and the canvas change depending on it should not ship in a single step | Passed | Passed |
| 7 | Prevention includes an automated pre-ship check comparing live columns and choice members with the app's cached metadata | Passed | Failed |
| 8 | Verification: perform the task in the published app on a fresh (uncached) player and confirm the written row in Dataverse | Passed | Failed |
2. A flow that writes to its own trigger table
9/9 and 9/9 with the skill 8/9 and 8/9 without
Write a solution-ready cloud flow: when a request is submitted, email and Teams-message the approver, then lock the same row.
With the skill
Without the skill
With the skill
Re-read the row before acting, guarded the lock on the value it changes, and passed the bundled flow linter in both runs.
Without the skill
Also loop-safe and lint-clean, but read everything from the trigger payload instead of re-reading the row: the one check it missed, in both runs.
All 9 checks for this test
| # | Check | With | Without |
|---|---|---|---|
| 1 | flow.json parses as JSON and has properties.definition with triggers and actions | Passed | Passed |
| 2 | Trigger is the Dataverse webhook on cr4f2_request with subscriptionRequest/message = 3 (Update) | Passed | Passed |
| 3 | Every connection reference uses runtimeSource embedded (not invoker) | Passed | Passed |
| 4 | The write of cr4f2_locked is guarded by a condition that reads cr4f2_locked (sentinel), so the self-write cannot loop | Passed | Passed |
| 5 | lint-flows.mjs reports no errors on flow.json | Passed | Passed |
| 6 | The row is re-read (GetItem/ListRecords) rather than relying only on the trigger payload | Passed | Failed |
| 7 | Email and Teams are not chained so that one runs after the other has Failed | Passed | Passed |
| 8 | notes.md covers the self-trigger loop risk and why the guard prevents it | Passed | Passed |
| 9 | notes.md covers deployment: connection references/consent and importing deactivated or turning on to validate | Passed | Passed |
3. Prove it in the published app
7/8 and 7/8 with the skill 3/8 and 4/8 without
Write a Playwright script that approves one pending request with a comment in the published canvas app and confirms it worked.
With the skill
Without the skill
With the skill
Kept the sign-in profile outside the project, defended against a stale cached player, warned that an admin run proves nothing about restrictions, and restored the row. Both runs made the Dataverse confirmation optional, ending as partial without it, which the tightened check fails.
Without the skill
Stored the sign-in state inside the project, had no defence against a stale cached player, and never queried Dataverse to confirm the write.
All 8 checks for this test
| # | Check | With | Without |
|---|---|---|---|
| 1 | verify-approve.mjs passes node --check (syntactically valid) | Passed | Passed |
| 2 | Uses a persistent browser profile / saved auth state kept outside the repo | Passed | Failed |
| 3 | Locates the app inside its iframe (frames / frameLocator) rather than only the top page | Passed | Passed |
| 4 | After filling the Comment box, blurs it (Tab) or otherwise commits the value before asserting/clicking | Passed | Passed |
| 5 | Handles the stale cached player (old-version banner, cache disabled, or build-stamp check) | Passed | Failed |
| 6 | Confirms the effect in Dataverse by DEFAULT: the script queries the row through the Web API on every run (not behind an optional flag or a commented-out block) and fails when the row was not updated | Failed | Failed |
| 7 | README states that an admin/owner session does not prove restrictions (non-admin needed) | Passed | Failed |
| 8 | README or script addresses that the approval writes production data and how to restore/record it | Passed | 1 of 2 |
4. Search and a 3,500-row picker in .pa.yaml
4/9 and 6/9 with the skill 2/9 and 2/9 without
Add a search box and a Dataverse vendor picker (about 3,500 rows) to an existing canvas screen kept as .pa.yaml, both filtering the gallery.
With the skill
Without the skill
With the skill
Fixed the compile-breaking text and passed the bundled .pa.yaml check in both runs, but both kept filtering the whole-table collection the screen already used, and neither notes file explained that Sum and lookup-column filters do not delegate. This test scored lower than on 0.1.0.
Without the skill
Left the original "Total open: " text in place, which breaks the whole-app compile, kept filtering the collection, and used the classic text input properties.
All 9 checks for this test
| # | Check | With | Without |
|---|---|---|---|
| 1 | Output .pa.yaml passes the bundled check-pa-yaml selftest rules (no colon-space in single-line values, no comments) | Passed | Failed |
| 2 | The pre-existing 'Total open: ' colon-space defect in lblTotal is fixed | Passed | Failed |
| 3 | Search box is a TextInput using Value/Placeholder (not Default/HintText) and is read as .Value | 1 of 2 | Failed |
| 4 | ComboBox includes a ComboBoxDataField child naming the display field | 1 of 2 | Failed |
| 5 | Vendor picker does not load all ~3,500 vendors into a collection or Items without addressing the row/search-page limit (it is server-side searched/filtered, or the limit is explicitly handled) | Passed | Passed |
| 6 | Gallery is filtered via a delegable query against the source rather than filtering a whole-table collection | Failed | Failed |
| 7 | notes.md mentions delegation and the 500/2,000-row limit | Passed | Passed |
| 8 | notes.md says Sum (an aggregate) over the source does not delegate, as a separate point from lookups | Failed | Failed |
| 9 | notes.md says filtering on a lookup's related column (e.g. Vendor.'Vendor Name') is not delegable, and the fix filters on the lookup itself or the vendor id | Failed | Failed |
5. Two flows that start each other
6/6 and 6/6 with the skill 6/6 and 6/6 without
Review two cloud flows before import: one recalculates an order when a line changes, the other copies the order's discount onto every line. Each has its own condition.
With the skill
Without the skill
With the skill
Found the cycle, ran the bundled linter on the originals (three errors) and on the fixes (none), and added filtering attributes so neither write can start the other flow.
Without the skill
Also found the cycle and fixed it with filtering attributes. The unaided model handles a loop this explicit, so this test guards against regression rather than separating the two.
All 6 checks for this test
| # | Check | With | Without |
|---|---|---|---|
| 1 | review.md identifies a loop BETWEEN the two flows (each one's write starts the other), naming both flows | Passed | Passed |
| 2 | review.md explains why each flow's own condition does not stop it (the If runs inside an already-started run, and/or the trigger has no filtering attributes, and/or utcNow() makes every write a change) | Passed | Passed |
| 3 | Both fixed flows set subscriptionRequest/filteringattributes on their update trigger | Passed | Passed |
| 4 | The fix breaks the cycle with a trigger-level condition or column filtering (not only an inner If) | Passed | Passed |
| 5 | lint-flows.mjs run over BOTH fixed flows together reports no errors (no trigger-cycle, no update-trigger-unfiltered) | Passed | Passed |
| 6 | Both fixed files still parse as JSON with properties.definition, and keep the original purpose (total recalculated, discount copied) | Passed | Passed |
6. Messages only to the testers
7/7 and 7/7 with the skill 6/7 and 6/7 without
Write a cloud flow that emails a requester and copies their manager, where nobody outside a tester allowlist may receive anything until go-live.
With the skill
Without the skill
With the skill
Sends live only when a setting equals an exact word; To and Cc both resolve from the allowlist; an empty list sends nothing; test mail names its intended recipient; the notes prove it from run history.
Without the skill
Equally careful with recipients, but both runs triggered on create as well as update, which the check requires to be update only.
All 7 checks for this test
| # | Check | With | Without |
|---|---|---|---|
| 1 | flow.json parses and its trigger is an Update on cr4f2_leaverequest with filteringattributes including cr4f2_status | Passed | Failed |
| 2 | Live sending requires an explicit mode value (e.g. a NotifyMode setting equal to an exact word); anything else, including a missing or blank setting, is treated as test | Passed | Passed |
| 3 | In test mode BOTH the To and the Cc resolve only from the allowlist (no expression path addresses cr4f2_requesteremail or cr4f2_manageremail unless live) | Passed | Passed |
| 4 | An empty or missing allowlist in test mode sends nothing (it does not fall back to the real recipients) | Passed | Passed |
| 5 | Test-mode messages are marked (subject prefix or banner naming the intended recipient) | Passed | Passed |
| 6 | lint-flows.mjs reports no errors on flow.json | Passed | Passed |
| 7 | notes.md says how to prove nobody else was emailed from evidence (e.g. reading the run history's actual recipients), not only from reading the source | Passed | Passed |
7. Long text in a gallery
5/6 and 6/6 with the skill 4/6 and 4/6 without
Users say descriptions, titles and names are cut off in a list. Fix the screen with the column lengths given so nothing is unreadably clipped and the full text stays reachable.
With the skill
Without the skill
With the skill
Clamped long values with an ellipsis and kept the full text reachable; the bundled format checker found no overflow in either run. One run left comments in the .pa.yaml, which breaks the compile.
Without the skill
Kept the full text reachable with a show-more toggle, but both runs put colon-space text in a single-line value, which breaks the compile, and one run still had labels the checker says will clip.
All 6 checks for this test
| # | Check | With | Without |
|---|---|---|---|
| 1 | The output .pa.yaml passes the bundled check-pa-yaml rules (no colon-space in single-line values, no comments) | 1 of 2 | Failed |
| 2 | check-canvas-format.mjs with the provided column lengths reports no text-overflow error on the output | Passed | 1 of 2 |
| 3 | The full description remains reachable (tooltip, detail view or a flexible-height row), not just truncated | Passed | Passed |
| 4 | Clamped values show a visible truncation marker (e.g. an ellipsis) rather than a silent cut | Passed | Passed |
| 5 | No label inside the gallery row relies on Overflow.Scroll | Passed | Passed |
| 6 | notes.md explains the approach chosen for long text and its trade-off | Passed | 1 of 2 |
8. A list nobody asked to filter
5/7 and 6/7 with the skill 6/7 and 5/7 without
Add an equipment list screen (about 1,800 rows, two choice columns) with an Edit button per row. Nothing is said about filters.
With the skill
Without the skill
With the skill
Added Category and Status filters with an All state and a name search, delegable at this size. One run broke the .pa.yaml check while reporting it clean, and neither run invited the user to change the defaults.
Without the skill
Added the same filters unprompted, but both runs left comments or colon-space text that breaks the compile. On this test the skill made no difference.
All 7 checks for this test
| # | Check | With | Without |
|---|---|---|---|
| 1 | The gallery can be narrowed by Category (a dropdown or similar control feeding Items) | Passed | Passed |
| 2 | The gallery can be narrowed by Status | Passed | Passed |
| 3 | There is a search on Name feeding the gallery's Items | Passed | Passed |
| 4 | The filters offer an 'all' state (blank or an explicit All) so the unfiltered list is reachable | Passed | Passed |
| 5 | The Items query stays delegable for ~1,800 rows (Filter on the source, no whole-table collection), or the limit is explicitly handled | Passed | Passed |
| 6 | The output passes the bundled check-pa-yaml rules | 1 of 2 | Failed |
| 7 | notes.md states which filters/sort were chosen as defaults and invites changes (or asks which ones are wanted) | Failed | 1 of 2 |
9. Theme before the first screen
7/7 and 6/7 with the skill 5/7 and 5/7 without
Before building a room-booking app, write the plan: what is needed from the user and in what order the app is built.
With the skill
Without the skill
With the skill
Asked for the palette, colours to avoid, fonts, logo and icon rules before any screen, defined them once as tokens, and planned verification in the published app.
Without the skill
Also asked for the theme first, but planned verification only in Studio and user testing, and never asked what each role should see first.
All 7 checks for this test
| # | Check | With | Without |
|---|---|---|---|
| 1 | plan.md asks for the brand/theme inputs (at least three of: palette, colours to avoid, fonts, logo or imagery, iconography or symbols, light/dark) before any screen is built | Passed | Passed |
| 2 | The build order puts theme definition before the first screen | Passed | Passed |
| 3 | Colours and fonts are defined once (theme tokens, named formulas or variables) and referenced by screens, not literal per control | Passed | Passed |
| 4 | plan.md asks what the landing page should show per role or audience | 1 of 2 | Failed |
| 5 | plan.md asks which filters, search or grouping the room list needs (or states a default) | Passed | Passed |
| 6 | The build order creates the Dataverse schema before the screens that bind it | Passed | Passed |
| 7 | plan.md includes verifying in the published app, not only in Studio | Passed | Failed |
10. The documentation set
7/7 and 7/7 with the skill 6/7 and 6/7 without
An expense approval app goes to pilot next week. Plan the handover documentation and outline each document.
With the skill
Without the skill
With the skill
Planned user, approver, administrator and developer guides with per-role screenshots from the published app using test data, and a checklist that every screen and role is covered.
Without the skill
Planned the same set and the same test-data rule, but nothing confirms that every screen and role is covered.
All 7 checks for this test
| # | Check | With | Without |
|---|---|---|---|
| 1 | Proposes separate documents per audience: at least the submitting user, the approver/manager, the administrator, and a developer/platform guide | Passed | Passed |
| 2 | The developer/platform guide outline covers environment/solution, Dataverse tables, security roles, connections or connection references, and each flow (trigger and what it writes/sends) | Passed | Passed |
| 3 | Screenshots are taken per role in the published app | Passed | Passed |
| 4 | Screenshots use test data, not real people's data (or the guides are kept out of public git for that reason) | Passed | Passed |
| 5 | States that statements must come from the app source or the live environment (button names as the controls say), not memory | Passed | Passed |
| 6 | Documents carry a version and issue date (or are tied to a release/build) | Passed | Passed |
| 7 | Includes a check that every screen and role is covered (an inventory or checklist) | Passed | Failed |
What the skill still misses
Four checks failed in every run, with the skill and without. They are the work before 1.0.
- Test 3, check 6: the verification script makes the Dataverse confirmation optional and ends as partial without it. The check asks for it on every run, failing when the row was not updated.
- Test 4, checks 6, 8 and 9: both runs kept filtering the whole-table collection the screen already loaded, and neither explained that
Sumover the source and a filter on a lookup's related column do not delegate. On 0.1.0 the skill moved the gallery to a delegable query; this is a regression to fix. - Test 8, check 7: the skill states the filters it chose by default but does not invite the user to change them. The baseline did in one run.
- The format checker resolves
ThisItem.Columnagainst the given column lengths but notGallery.Selected.Column, so it guessed a length for a detail label in test 7. The test's schema now names that column; the checker should resolve it.
What it costs
The skill's reference files are read only when a task needs them, but they are read, and most of the extra tokens are those files read from cache. Expect a slower, more expensive answer in exchange for the checks above.
- Checks passed
- 90% with, 66% without
- Mean tokens per task
- 507k with, 173k without; output 14k against 12k
- Mean time per task
- 142 s with, 116 s without
History
The two results are not comparable point for point: three 0.1.0 checks were tightened (one split into three), and six tests were added that the unaided model already handles fairly well.
| Version | Date | Tests and runs | With the skill | Without |
|---|---|---|---|---|
| 0.1.0 | 2026-10-01 | 4 tests, 32 checks, 1 run each | 32/32 (100%) | 14/32 (44%) |
| 0.5.1 | 2026-10-02 | 10 tests, 74 checks, 2 runs each | 133/148 (90%) | 98/148 (66%) |
| 0.5.1 | 2026-10-02 | tests 1 to 4 only (the 0.1.0 set, tightened) | 58/68 (85%) | 33/68 (49%) |
| 0.5.1 | 2026-10-02 | tests 5 to 10 only (new) | 75/80 (94%) | 65/80 (81%) |
Does it load when it should?
A skill that never loads helps no one, and one that loads for the wrong work wastes tokens. Twenty requests were written for 0.1.0: ten that need the skill and ten near misses that share its vocabulary but not its job, such as a Power BI DAX measure, a Dynamics 365 C# plug-in, a Power Automate Desktop flow, a Logic Apps workflow in Bicep and a React client for the Dataverse Web API.
- Held-out requests
- 8 of 8 correct
- Training requests
- 11 of 12 correct
- False triggers
- 0 in every round
The skill's description has not changed since 0.1.0, so this evaluation was not re-run. In the 0.5.1 task runs the skill loaded in all 20 runs where it was installed.
How it was measured, and its limits
- Same model, same prompt, separate sessions. Each run was a fresh headless Claude Code session in its own folder. For the runs with the skill it was installed as a project skill; the runs without it had none. Both got the same instruction to stay offline.
- Machine checks where possible. Flow definitions went through the bundled flow linter (both flows of test 5 together),
.pa.yamlfiles through the bundled compile-fault check and the long-text checker with the column lengths given, and scripts throughnode --check. A grader model judged the rest without knowing which configuration produced the answer, and had to quote its evidence. - Two runs per configuration. With the skill, six of ten tests scored the same in both runs. The spread is shown per check as a ring. Two runs are still a thin variance estimate.
- 47 of 74 checks passed in every run of both configurations. They guard against regressions but do not separate the two. The gap comes from the 22 that favour the skill. The unaided model already catches an explicit two-flow loop (test 5) and adds list filters when it sees choice columns (test 8).
- Nothing ran against a live tenant. The checks judge the answers on paper. The skill's own rule is that only performing the task in the published app proves a change, and this evaluation could not do that.