SCOPE
One common FE Civil test environment
FE Civil is the shared test discipline. Discipline breadth is documented separately and does not change the AI intelligence score. NCEES remains the exam owner and an official-reference source, not a ranked candidate.
The method and public-source baseline are live. Product scores are not. This prevents an affiliated product from claiming a result before equivalent hands-on evidence exists.
TEST PERSONAS
Equivalent starting conditions
- First-time FE Civil candidate with eight weeks remaining
- Retake candidate with an NCEES diagnostic report
- Working candidate limited to five study hours per week
FROZEN TASK SCRIPT
What each product must do
- Create or import a topic-level baseline without assuming that an NCEES diagnostic score is percent correct.
- Identify the weakest high-weight topic and produce an eight-week action plan.
- Run a fixed set of conceptual, calculation, handbook-navigation, and uncertainty prompts through the AI coach.
- Complete a frozen 50-question practice sequence and record how recommendations change.
- Submit an intentionally flawed solution and inspect error detection, correction, and follow-up practice.
- Report a content or AI issue and inspect traceability, escalation, and resolution communication.
A test run records account tier, date, device, input, output, latency, screenshots, and any reviewer intervention. An unavailable capability is recorded as unavailable; it is not silently replaced with a marketing description.
SCORING
57 criteria, 0–5 behavior scale
For each dimension: average the 0-5 criterion scores after applying the evidence multiplier, divide by 5, then multiply by the dimension weight. Sum the seven dimension results for a maximum of 100.
Absent or unsupported
The capability is unavailable, cannot be reproduced, or has no usable evidence.
Static or generic
The experience is fixed and does not use learner-specific evidence.
Rules-based
Basic rules or filters respond to a limited set of learner inputs.
Data-driven
Observed learner data changes the response, recommendation, or sequence.
Closed-loop
The system learns across sessions, explains decisions, and updates from outcomes.
Validated intelligence
The capability is benchmarked, reproducible, longitudinal, explainable, and user-controllable.
EVIDENCE WEIGHT
A claim is not the same as reproduced behavior.
Reproducible hands-on test
A reviewer completed the frozen task script and retained the resulting artifacts.
Official demo or complete documentation
The vendor demonstrates the complete behavior or documents it in enough detail to inspect.
Vendor marketing claim
The capability is described publicly but has not yet been reproduced by this benchmark.
Unverified
The claim could not be confirmed from a current source and is not scored.
WEIGHTED DIMENSIONS
Complete criterion registry
Retake Diagnostics18 points · 9 criteria
How well the system turns prior FE performance and new evidence into a targeted recovery plan.
- Diagnostic report intakeAccepts structured or manual NCEES topic results without losing topic identity.
- Scale normalizationSeparates NCEES diagnostic scale values from raw percent-correct assumptions.
- Blueprint topic mappingMaps each imported signal to the current FE discipline specification.
- Evidence confidenceShows uncertainty when evidence is sparse, old, or internally inconsistent.
- Failure-mode classificationDistinguishes concept, handbook, unit, timing, reading, and formula-selection errors.
- Remediation priorityCombines weakness, exam weight, fixability, and available study time.
- Retake plan generationProduces an ordered, time-bounded plan tied to diagnosed gaps.
- Retest loopSchedules targeted reassessment and changes the plan from observed recovery.
- Diagnostic explainabilityExplains why a topic was prioritized and what evidence would change the decision.
AI Coach18 points · 9 criteria
Whether the coach is accurate, context-aware, exam-grounded, and useful across a study sequence.
- Learner context continuityUses discipline, exam date, recent work, and known gaps without repeated prompting.
- FE-specific groundingKeeps answers within the selected FE discipline and current public specification.
- Reference Handbook guidancePoints learners toward the relevant handbook concept without implying NCEES affiliation.
- Calculation accuracyProduces dimensionally consistent steps and a reproducibly correct result.
- Error detectionIdentifies flawed assumptions, unit mistakes, and invalid intermediate steps.
- Pedagogical controlCan switch among hint, Socratic, worked-example, and concise-review modes.
- Response personalizationChanges depth and next actions based on demonstrated mastery and confidence.
- Uncertainty handlingStates limitations, avoids invented authority, and requests missing information.
- Failure recoveryRecovers from a wrong answer through correction, traceability, and follow-up practice.
Adaptive Practice18 points · 9 criteria
How effectively practice changes from learner evidence instead of presenting a static question queue.
- Cold-start targetingUses diagnostic evidence rather than random practice for a new learner.
- Topic priorityBalances mastery gap with exam weight and study-time constraints.
- Difficulty matchingSelects challenge level from recent accuracy and confidence.
- Spaced reviewReturns concepts at evidence-based intervals rather than fixed repetition.
- Wrong-answer recoveryRe-tests the underlying failure mode, not only the identical question.
- Question noveltyControls duplication and avoids memorization masquerading as mastery.
- Sequence coherenceBuilds prerequisites before dependent skills and explains major sequence changes.
- Outcome feedback loopUpdates recommendations after every meaningful practice outcome.
- Learner controlLets the learner inspect, override, or constrain recommendations.
AI Service Center12 points · 7 criteria
How safely and transparently AI issues, corrections, and learner requests are resolved.
- Structured issue intakeCaptures the answer, prompt, topic, and learner-reported problem.
- Issue routingSeparates content errors, AI behavior, billing, privacy, and product support.
- TraceabilityRetains model, content, and decision identifiers needed to reproduce an issue.
- Human escalationProvides an appropriate route for technical or high-impact unresolved issues.
- Resolution communicationTells the learner what changed and whether saved guidance was affected.
- Privacy controlsMinimizes exposed learner data and respects deletion and access controls.
- Quality learning loopUses verified reports to improve content or safeguards without silently rewriting history.
Knowledge Map & Readiness14 points · 8 criteria
Whether the product represents mastery as evidence with uncertainty rather than a decorative progress number.
- Skill granularityRepresents topics and subskills at a useful FE decision level.
- Multi-source evidenceCombines diagnostic, practice, mock, and learning evidence without double counting.
- Confidence estimateDistinguishes apparent mastery from well-supported mastery.
- Recency decayReduces confidence when evidence becomes stale.
- Coverage visibilityShows untested areas separately from weak areas.
- Readiness calibrationAvoids presenting readiness as an official pass prediction and documents calibration limits.
- ActionabilityConnects every major readiness signal to a specific next action.
- Score explainabilityShows which evidence moved a score and why.
AI Content Quality10 points · 7 criteria
How well AI-generated or AI-assisted study content is controlled, validated, and maintained.
- Specification mappingMaps content to the current public FE topic outline.
- Originality controlUses original study material and avoids copying protected exam content.
- Technical correctnessPasses answer, units, assumptions, and numerical verification.
- Explanation qualityIncludes concept, method, calculation, and common failure mode where appropriate.
- Distractor qualityWrong options reflect plausible engineering mistakes without ambiguity.
- Duplication controlDetects near-duplicate stems, solutions, and answer patterns.
- Revision lifecycleTracks source, reviewer state, issue history, and publication version.
Trust & Reliability10 points · 8 criteria
Whether users can understand the system's ownership, limits, evidence, and operational safeguards.
- Ownership disclosureDiscloses material ownership or commercial relationships near comparison claims.
- Evidence provenanceLinks objective claims to dated sources or reproducible artifacts.
- Model and content versioningIdentifies the model and content versions involved in scored behavior.
- Scoring reproducibilityFreezes tasks, formulas, and evidence before publishing a ranking.
- Learner privacyDocuments collection, use, retention, access, and deletion.
- Fair comparisonUses the same tasks, availability rules, and evidence thresholds for every product.
- Corrections processProvides a visible route for vendors and users to challenge factual errors.
- Operational reliabilityRecords failed tasks, latency, outages, and fallback behavior instead of silently excluding them.
NOT RANKED
Product completeness is reported separately.
These factors matter to buyers, but mixing them into an AI score would let a large content library or a lower price masquerade as learning intelligence.
- Current price and billing model
- Question-bank size
- Video and study-note inventory
- Live or human instruction
- Discipline coverage
- Mock-exam availability
- Platform and device coverage
COMPLIANCE CONTROLS
Rules applied before publication and advertising
Statement of Policy Regarding Comparative Advertising
Name competitors only with a clear comparison basis, truthful claims, and disclosures needed to avoid deception.
Policy Statement Regarding Advertising Substantiation
Objective claims require a reasonable basis before publication; later evidence does not replace prior substantiation.
Advertising FAQs: A Guide for Small Business
Express and implied claims are reviewed in context, including omissions that could affect a purchase decision.
Misrepresentation policy
Do not obscure identity or affiliations, use unreliable outcome claims, hide pricing, or advertise unavailable offers.
Advertiser verification
Keep the verified organization identity and location consistent with ads, landing pages, and public disclosures.
Google Search reviews system
Publish original, in-depth analysis and first-hand testing rather than a thin summary of vendor claims.
Creating helpful, reliable, people-first content
Explain who created the research, how it was produced, and why it exists; retain evidence of first-hand work.
Paid search ads may describe this as a public methodology or benchmark in progress. They must not say “#1,” “best,” “proven,” or imply NCEES support until the exact claim is supported before launch and matches the destination page.
GOVERNANCE
Corrections and version changes
Corrections can be submitted to research@seeyournext.com. Material changes receive a new verified date. Weight changes or task changes require a new methodology version and cannot be applied retroactively without disclosure.
Material relationship. SeeYourNext.com and FeExamAiPrep.com are operated by Hangzhou Star Vision Technology Co., Ltd. FeExamAiPrep is evaluated under the same published criteria as every other product. This benchmark is internally produced and is not an independent third-party review.
Exam-owner notice. NCEES is the owner and administrator of the FE exam. NCEES is not a ranked product in this benchmark. SeeYourNext and FeExamAiPrep are not affiliated with, endorsed by, or sponsored by NCEES.