Skip to main content
Ground Control | Tier 1
August 26, 2026
Question

I am trying to run the Agentic Potential Scoring tool and giving some inputs, the same is stored in my google sheetas well, but ater executing the tool it changes to some other keys for the IDs and it does not work. Where am I going wrong?

  • August 26, 2026
  • 7 replies
  • 50 views
My input for sOpportunities


[
  {
    "id": "10293847",
    "score": 43,
    "reasoning": "Impact 8 (240h/yr), Vol 3 (<100/mo), Stab 8 (doc SOP), Exc 5 (assume 15%), Feas 6 (email/Excel), Risk 5 (assume mod), Align 3 (no goal), Ready 5 (owner yes, no sponsor). Missing: exception%, data, sponsor, goal. Park/Decline <50."
  },
  {
    "id": "11847562",
    "score": 71,
    "reasoning": "Impact 16 (1200h/yr), Vol 8 (1000+/mo), Stab 11 (doc/stable), Exc 7 (assume 15%), Feas 10 (web apps), Risk 7 (assume mod), Align 5 (indirect BU), Ready 7 (owner+tentative). Quick Win ≥70. Assumptions: exception%, exact vol."
  },
  {
    "id": "12938475",
    "score": 58,
    "reasoning": "Impact 12 (720h/yr), Vol 7 (100-999/mo), Stab 5 (outdated SOP), Exc 5 (assume 15%), Feas 8 (CSV/email), Risk 5 (assume mod), Align 8 (cost reduction goal), Ready 8 (owner+support). Investigate. Missing: exception%, data."
  },
  {
    "id": "13029384",
    "score": 35,
    "reasoning": "Impact 4 (180h/yr), Vol 2 (<10/mo), Stab 3 (no SOP), Exc 5 (assume 15%), Feas 5 (assume common apps), Risk 5 (assume mod), Align 3 (no goal), Ready 8 (owner+sponsor named). Park <50. Low vol/impact."
  }
]
This is what I am getting in the scoredOpportunities variable, and this has different IDs,why and how to fix this?

    7 replies

    Aaron.Gleason
    Automation Anywhere Team
    Automation Anywhere Team
    August 26, 2026

    @nehavenkateshmurthy Interesting…  🤔

    Since the input IDs don’t match the output IDs, the LLM is inventing IDs. If you could include your System and User Prompts, that would help us figure this out. For example, we would need to check whether your prompts mention keeping the same ID from each input opportunity or if there is just a generic “return this JSON” message. The generic message tends to cause LLMs to fabricate IDs.

    Ground Control | Tier 1
    August 26, 2026

    I am just using user prompt in AI skill like below

    # Context You score automation opportunity submissions for an automation Center of Excellence operating a Start-stage pipeline. Input is a JSON array. Each element has: - `id` — unique 8-digit string. - `datetime` — submission timestamp. - `opportunity` — a plain string of newline-separated `Label: value` pairs captured by the intake form. Values are unverified self-reports from the requestor. ===== BEGIN SCORING CRITERIA (verbatim — replace this block if the criteria document changes) ===== Red Flags (No-Go until resolved) - Underlying platform/application replacement or policy prohibitions - Unavailable credentials/test data or unresolved licensing - Mission-critical flow where automation failure is unacceptable without robust controls - No accountable business owner/sponsor Triage Gate (Pass/Fail) — any No = Fail: Owner, Stable, Access, No decommission, Data manageable, Sponsor ready. How to score (fast + fair) 1. Capture the best available estimates (volume, AHT, exceptions). 2. Use the anchor that most closely fits and note assumptions. 3. If split between two anchors, pick the lower and add a comment. 4. If a Red Flag is present, pause scoring and remediate first. 5. Re-score after discovery if key inputs change. 1) Business Impact — 0–20 Purpose: Size potential value (time/$$/risk/CX). Evidence: Volume/month; AHT (minutes); % time saved; error/SLA data. Quick calc (hours/yr): (Volume/month × AHT × % saved × 12) ÷ 60. - 0–5 (Low): <500 hours/year or <$$50k; limited risk/CX lift. - 10–15 (Med): 500–1,000 hours/year or $$50–$$100k; noticeable risk/CX benefit. - 16–20 (High): >1,000 hours/year or >$$100k; material CX/compliance gain. Common pitfalls: Inflated % saved; double-counting downstream benefits. 2) Volume & Frequency — 0–10 Purpose: Ensure enough repetitions to matter. Evidence: Transactions/month; arrival pattern; peak windows. - 0–3 (Low): <500/month; ad-hoc; minimal batching. - 4–7 (Med): 500–5,000/month; routine cadence; some peaks. - 8–10 (High): >5,000/month or daily peaks causing SLA risk. Common pitfalls: Counting records instead of actual handled transactions. 3) Process Stability — 0–15 Purpose: Predictability of steps/inputs/rules. Evidence: SOPs; change cadence; variants by segment/region. - 0–4 (Low): No SOP; steps vary by person; upstream change ~bi-weekly. - 5–10 (Med): Documented SOP; 1–2 variants; minor quarterly changes. - 11–15 (High): Standardized; governed change management; few stable variants. Common pitfalls: Hidden "tribal" steps; frequent policy updates not captured. 4) Exception Rate — 0–10 (lower is better) Purpose: Degree of straight-through vs. manual judgement. Evidence: % work off happy path; rationale categories. - 0–3 (Low score): >20% exceptions; judgement-heavy. - 4–7 (Mid): 10–20% with recognizable patterns. - 8–10 (High score): <10% exceptions; rule-based, routable. Common pitfalls: Treating rework loops as exceptions (count once per case). 5) Technical Feasibility — 0–15 Purpose: Fit with current platforms/connectivity. Evidence: Apps (web/desktop/API), access model, controls, automatable artefacts. - 0–4 (Low): Bot-hostile UI only; unstable remoting; unknown endpoints. - 5–10 (Med): Common apps (Web/Excel/Outlook/SharePoint); simple integrations. - 11–15 (High): APIs/structured data; supported packages; low environment friction. Common pitfalls: Underestimating multi-factor auth, VDI latency, pop-ups. 6) Risk & Data Suitability — 0–10 (higher = safer) Purpose: Data classification + control adequacy. Evidence: Data types (PII/PCI/PHI), access model, logging/audit options. - 0–3 (Low): Sensitive data with inadequate guardrails; policy red flags. - 4–7 (Med): Moderate sensitivity but controllable via vault/RBAC; audit feasible. - 8–10 (High): Low sensitivity; straightforward audit trail; minimal policy burden. Common pitfalls: Non-prod data unavailability; copying data outside approved zones. 7) Strategic Alignment — 0–10 Purpose: Link to current organizational goals, roadmaps and portfolio themes. Evidence: Named goals and KPIs; BU priority; linkage to program themes. - 0–3 (Low): No tie to OKRs; one-off request. - 4–7 (Med): Indirect link; BU-level priority; enabling capability. - 8–10 (High): Directly advances top enterprise/BU objective this quarter. Common pitfalls: "Pet project" bias without measurable outcome. 8) Stakeholder Readiness — 0–10 Purpose: Willingness/ability to partner and adopt. Evidence: Named sponsor/owner; SME time; adoption/change plan. - 0–3 (Low): No sponsor; change resistance likely; SME access uncertain. - 4–7 (Med): Owner named; tentative support; limited SME time. - 8–10 (High): Committed sponsor; SME hours booked; adoption owner identified. Common pitfalls: Shadow sponsors vs. accountable leader; over-promised SME time. Total Score (0–100). Quick-Win Checklist & Decision Guide - Quick Win (green-light to backlog): Score ≥ 70; no triage fails and no red flags; ≤ 2 primary systems, ≤ 4 weeks build estimated, ≤ 10% exception rate. - Investigate Further: Score 50–69 or 1 remediation needed (e.g., secure data access). Time-box discovery (≤ 2 weeks) and re-score. - Park / Decline: Score < 50 or any non-remediable Red Flag. Tie-breakers & edge cases - Similar totals: Prefer higher impact × lower effort (feasibility + exceptions as effort proxy). - High impact but unstable: Time-box discovery for stabilization plan; don't push to Quick Win yet. - Low impact but high alignment: Bundle into a themed release or deprioritize. Minimum evidence checklist (per submission): Volume/month, average handling time, exception %, systems (top 3), data category, named owner/sponsor, organizational goal link. ===== END SCORING CRITERIA ===== ===== BEGIN FIELD MAPPING (replace this block if the intake form changes) ===== - `Name` → Stakeholder Readiness. Requestor identity only; not evidence of a sponsor. - `Email` → no factor. Identity and deduplication. - `Date` → no factor. Submission date. - `What's the work you want help with?` → Business Impact (scope of work) and Technical Feasibility (systems implied by the description). - `Why is it painful today? Who feels it?` → Business Impact (pain, error/SLA signals, population affected). Also Strategic Alignment, but only when a named goal, KPI, or OKR appears. - `How much time does it take to do this work, once (minutes)?` → Business Impact. This is AHT in the hours/yr calc. - `How often does this happen?` → Volume & Frequency, and the volume input to Business Impact. Form bands are narrower than the criteria bands: `<10 times/month`, `10–99/month`, and `100–999/month` are all <500/month. `1,000+/month` is unbounded — score it Med unless the text evidences >5,000/month, and flag the assumption. - `What kicks it off and what's the input?` → Technical Feasibility (trigger and input artefact). `System event` and `Online form` indicate structured/API input; `File (CSV/PDF/Excel)` and `Email arrives` indicate common-app integration; `Scheduled/time-based` also informs arrival pattern for Volume & Frequency; `Other` alone is not evidence. - `Is the process well documented?` → Process Stability. `Yes – I could send them to you now` = documented SOP; `Yes – But they're very outdated` = SOP exists but unreliable; `I don't know` and `No` = no usable SOP. - `Do you manage the process?` → Stakeholder Readiness. `Yes` = the requestor is the accountable owner; `No` = owner unidentified, which alone does not evidence the "No accountable business owner/sponsor" red flag. No field collects: exception rate, data sensitivity, strategic alignment, the systems list for Technical Feasibility, % time saved for Business Impact, or sponsor/SME commitment. Score these from free-text evidence where it exists. ===== END FIELD MAPPING ===== # Objective Score each element of the input array independently. 1. Parse `opportunity` into label-value pairs and route each to its factor using the field mapping. 2. Score all 8 factors from evidence present in that submission only. Do not infer volumes, savings, sponsorship, data sensitivity, or system landscape that the text does not state. 3. Where a factor has no supporting evidence, assign the midpoint of its Med band rounded down, and name the missing input as an assumption in `reasoning`. 4. Sum the 8 factors into `score`. 5. Apply the Decision Guide and name the resulting band in `reasoning`. 6. If the submission explicitly evidences a Red Flag, prefix `reasoning` with `RED FLAG: <flag>` and set the band to Park / Decline. Never assert a red flag from absent data. 7. In `reasoning`, cite the dominant scoring drivers and every assumption you made. # Response Format Return a JSON array with one object per input element, in input order, and nothing else — no markdown, no backticks, no surrounding text. {   "id": "abc-123",   "score": 90,   "reasoning": "Why it got this score" } `reasoning` must not exceed 255 characters. The opportunities to assess are: $sOpportunities$  

    Aaron.Gleason
    Automation Anywhere Team
    Automation Anywhere Team
    August 26, 2026

    @nehavenkateshmurthy Thank you. Who better to ask about LLM issues than an LLM? 

    Claude says that this output seems to have nothing to do with the prompt’s input. Claude thinks these results are fabricated (by the LLM, not you) or leftover results from a previous run.

    If that’s not the case, the LLM you’re using is seriously having issues. I might swap over to another LLM and see if your results are better. 

    Edit: You can also see if you can find any of those IDs anywhere in your input dataset.

    Ground Control | Tier 1
    August 26, 2026

    Okay let me try that. Thanks!

    Ground Control | Tier 1
    August 26, 2026

    I improved the prompt and it is working now. Thanks for prompt response. ​@Aaron.Gleason 

    Aaron.Gleason
    Automation Anywhere Team
    Automation Anywhere Team
    August 26, 2026

    @nehavenkateshmurthy Please, PLEASE tell me that was an intentional pun.  🤣🤣🤣

    Ground Control | Tier 1
    August 26, 2026

    Haha honestly, it was!! 🤣