Back to Home

Validation Study

1. Executive Summary

Athir is a standardised, evidence-linked behavioural assessment for executive selection and development. Candidates respond to simulated executive scenarios; the Athir Behavioural Intelligence Engine scores observed behaviour against calibrated competency indicators; every score is linked to timestamped behavioural evidence.

This Validation Study describes the evidence supporting the interpretation and use of Athir assessment scores. It is written to be read critically by occupational psychologists, HR leaders, procurement teams, legal counsel and analysts. It distinguishes clearly between: what Athir has documented internally; what has been tested only on synthetic data; what is supported by external scientific literature; and what has not yet been tested.

Athir does not currently publish a criterion-related validity coefficient, reliability coefficients, or a factor-analytic study of its competency structure. Each will be published only when a documented evidence package exists that allows an external reviewer to reproduce the calculation.

What this document supports today
Documented internally

A documented assessment architecture: eight-stage simulations, standardised scoring rubrics, deterministic score mapping, and a traceable evidence trail from every score to the candidate interaction that produced it.

Synthetic testing

An internal fairness audit using 2,847 synthetic responses: preliminary, pre-deployment algorithmic fairness testing (see the Bias & Fairness Audit).

External literature

Methodological alignment with recognised professional frameworks, cited in the References section. External literature supports methodology choices; it is never evidence of Athir-specific performance.

Not yet established

Criterion-related validity, reliability, internal structure and real-world fairness: not yet established. The evidence maturity statement (Section 20) states Athir's current position for each area.

2. Intended Purpose and Use

Athir is designed to inform executive selection, promotion and development decisions. It is a decision-support instrument: it is designed to inform human decision-making, not to autonomously appoint or reject candidates. No candidate should be selected or rejected solely because of an Athir output; appropriately trained humans retain decision authority, and reviewers should be able to challenge outputs.

Intended population: executives and senior leaders, assessed in English, in enterprise and business-unit contexts corresponding to the four Target Benchmarks described in Section 4.

Intended interpretation: Athir assesses general executive capability as exhibited under standardised simulated conditions, contextualised by the expected organisational level (Target Benchmark) and the primary business domain (Business Function). It does not claim to predict role-specific fit, and it does not claim certainty about future executive performance.

3. Assessment Architecture

An Athir assessment is an eight-stage behavioural simulation. At each stage the candidate receives staged business communications and interacts with simulated stakeholders: reading or replying to messages, making decisions, delegating, allocating budgets and setting priorities under time pressure. The platform records the full interaction stream: every response, the time taken, the decisions selected, and the scenario that prompted them.

The Behavioural Intelligence Engine evaluates recorded behaviour against documented scoring rubrics for the relevant competency. The engine applies the same rubric version to every candidate in a given assessment configuration. The assessment record preserves scenario identifiers, stage timing, and model and framework versions, so that any score can be traced back through the chain described in Section 12.

4. Executive Assessment Context

Every assessment is configured with two context variables:

  • Target Benchmark: the organisational level at which the executive is expected to operate (Enterprise Executive, Business Unit Executive, Functional Executive, Operational Leader).
  • Business Function: the primary business domain the executive leads (for example Strategy, Finance, People, Commercial, Operations, Supply Chain, Technology, Product, Governance, Corporate Services).

Together these define the Executive Assessment Context. Context shapes scenario selection, business terminology and the level of expectation applied when interpreting scores. The same scoring criteria and rubric versions are applied to every candidate regardless of demographic attributes; whether score outcomes are equivalent across demographic groups in real populations is an empirical question, and testing to date is synthetic and preliminary (Section 11).

5. Competency Framework

The Athir framework comprises eight competencies: Strategic Thinking, Decision Quality, Leadership Influence, Execution Discipline, Risk Judgement, Communication Clarity, Financial Acumen, and Enterprise Orientation. Each has a documented definition and a set of behavioural indicators used in scoring. The framework was developed with reference to the executive selection and assessment literature and to the practical demands of executive roles.

Formal content-validity evidence: a structured job analysis, an independent subject-matter-expert panel with documented composition and ratings, and a documented competency-to-indicator-to-scenario mapping, is being compiled. Until it is complete, Athir does not claim formal content validity, and the eight competencies are not presented as universal determinants of executive success.

What Athir assesses is a defined combination: (1) general executive capability exhibited under simulated conditions, contextualised by (2) level-specific expectations (Target Benchmark) and (3) functional context (Business Function). Athir does not claim role-specific fit prediction.

6. Scenario and Behavioural Evidence Methodology

Scenarios are drawn from a versioned scenario bank, allocated according to the Executive Assessment Context, and presented in a fixed stage structure so that candidates assessed for the same configuration face comparable demands. Every candidate interaction is logged with its type, timestamp, response time, the scenario that prompted it, and the behavioural indicator evidence extracted from it.

A typical assessment generates a large volume of recorded material: raw interaction events, the text of each response, timing data, decisions selected, and the behavioural indicators derived from them. A numeric “data point” count is not published in this document because the definition of a single data point (raw event, linguistic feature, behavioural indicator, scored observation, or model-derived feature) and a documented counting method are being finalised. A large observation count is not a proxy for validity; its methodological purpose is evidential traceability and audit.

Where a stage does not separately evidence a competency, the scoring methodology records a neutral score with an explicit “not separately evidenced in this stage” note rather than inferring performance.

7. Scoring Methodology

Each recorded interaction is mapped to one of the eight competency pillars and evaluated against the behavioural indicators in the applicable rubric, producing a pillar score on a 0 to 10 scale. The composite score is the unweighted mean of the eight pillar scores.

The overall readiness rating (Not Ready, Development Needed, Ready with Support, Ready Now, Exceptional) is generated by fixed, documented score rules applied to the composite score. Athir reports readiness only. It issues no appointment, rejection or selection verdict: the appointment decision is made by the human decision-makers the assessment informs.

These mappings are deterministic and reproducible: the same score profile produces the same rating under the same rule version. They are design decisions, not empirically derived cut scores, and they carry no demonstrated criterion validity. Readiness ratings inform human decision-making; they are not predictions of appointment success.

8. Score Interpretation

  • Interpret scores relative to the selected Target Benchmark and captured Business Function, not in isolation.
  • A score represents observed behavioural performance under simulated executive conditions on the assessment date, not a fixed trait.
  • Interpret scores at band level. Until the standard error of measurement is published, small differences (for example, 7.1 versus 7.3) should not be treated as meaningfully distinguishable.
  • Where scores fall near a readiness threshold, treat the classification as sensitive to measurement uncertainty rather than decisive.
  • Read scores alongside the timestamped behavioural evidence, the readiness rating and the board recommendation in the report.

9. Evidence for Validity

9.1 Content evidence

The competency framework, scenario architecture and rubrics are documented, versioned, and mapped to the intended population and context. Formal content-validity evidence (job analysis, independent SME review) is in progress; Athir does not currently claim formal content validity (Section 5).

9.2 Internal structure

No factor-analytic evidence is currently published. A pre-registered study of the framework's internal structure (exploratory and confirmatory analysis with reported loadings and fit indices) is planned. The eight-pillar structure is a documented design model, not an empirically demonstrated factor structure.

9.3 Criterion-related evidence

Athir does not currently publish a criterion-related validity coefficient. Athir distinguishes carefully between concurrent validity (scores related to performance measured at approximately the same time) and predictive validity (assessment administered before subsequent performance is observed); retrospective or concurrent evidence will never be described as predictive validity.

When criterion-related studies are completed, results will be published in full form: per criterion and per study first, with design, sample, blinding, corrections and confidence intervals. No aggregate validity range will be published without an exact statistical meaning for the range.

10. Reliability and Measurement Precision

Athir does not currently publish reliability coefficients. A single coefficient across eight distinct constructs would not be appropriate evidence of quality, and for an automated engine, agreement with human experts and inter-rater reliability are different questions.

Planned reliability evidence, reported at the level scores are interpreted:

  • Per-competency internal consistency, or coefficient omega where assumptions warrant it, rather than a single across-construct alpha.
  • A test-retest study documenting sample, interval, whether parallel or identical scenario forms were used, and practice effects.
  • An engine repeatability study: score variance when identical responses are scored under the same rubric version.
  • The standard error of measurement and confidence intervals around individual scores, enabling uncertainty-aware interpretation near readiness thresholds.

A high test-retest coefficient, when published, will show score stability under the conditions studied; it will not by itself establish what underlying construct produces that stability, and Athir will not word it as if it did.

11. Fairness and Adverse-Impact Evidence

Athir's current fairness evidence is an internal audit conducted on 2,847 synthetic responses across five demographic dimensions: preliminary, pre-deployment algorithmic fairness testing. No statistically significant group differences were detected at α = .05. This finding does not demonstrate the absence of demographic bias in real candidate populations.

The audit, its methodology and its limitations are documented in full in the Bias & Fairness Audit. Real-world demographic fairness monitoring is planned from the 2027 audit cycle onward, subject to sufficient real-world data.

12. Explainability and Audit Trail

Athir describes its architecture as a traceable evidence architecture rather than complete explainability. A qualified reviewer can trace an assessment output through the competency score, behavioural indicator and scored observation to the candidate interaction, the scenario, the applicable scoring rubric, and the model and framework versions used.

Each assessment record preserves: timestamps for every interaction; the scenario identifiers and versions served; the competency framework version; the scoring rubric version; the Target Benchmark and Business Function; the assessment date; and the processing history. Model-assisted inference steps are not claimed to be fully reconstructable; where they are not, the evidence trail carries the rubric version and indicators applied rather than a full inference transcript.

Athir does not claim that this architecture makes assessment decisions objective in any absolute sense. It makes them standardised, evidence-linked and reviewable.

13. Appropriate Use

  • Informing executive selection, promotion and development decisions alongside interviews, references and the judgement of trained assessors.
  • Structuring evidence-based discussion among hiring stakeholders.
  • Identifying development priorities, informed by the evidence summary.
  • Comparing candidates assessed for the same Target Benchmark, while respecting the limitations described in this document.

14. Inappropriate and Unsupported Uses

  • Selecting or rejecting any candidate solely on the basis of Athir scores, ratings or recommendations.
  • Inferring future performance with certainty, or treating a score as a fixed trait rather than observed behaviour under simulated conditions on the assessment date.
  • Use outside the intended population (executives and senior leaders) or in languages for which validation has not been conducted.
  • Interpreting small score differences as meaningful; differences within the same readiness band should not drive decisions.
  • Claiming, on the basis of an Athir assessment, the absence of demographic bias in a client's wider hiring process.
  • Any use as a standalone legal-compliance instrument.

15. Known Limitations

Athir's validation evidence is subject to the following limitations:

  • No criterion-related validity coefficient is currently published; the relationship between Athir scores and subsequent workplace performance has not yet been established.
  • Reliability coefficients are not currently published; score precision (standard error of measurement) is not yet established.
  • The current fairness audit uses synthetic data and is preliminary; it cannot demonstrate fairness in deployed populations.
  • Disability and intersectional fairness have not been statistically tested.
  • Scenario content neutrality has not been formally reviewed; a structured Scenario Fairness Review is planned.
  • Validation to date is English-language, in executive populations; findings do not generalise to other languages, cultures or populations without further evidence.
  • Readiness thresholds and the equal-weighted composite are design decisions, not empirically derived values.
  • Scoring is compensatory; a high pillar score can offset a low one, and no minimum competency threshold currently exists.
  • All current analyses were conducted internally by Athir. Nothing in Athir's documentation has been independently audited.

16. Generalisability

Athir's validation evidence is developed primarily for English-language executive contexts. This is a scope limitation, not a global claim: fairness and validity findings must not be generalised to non-English assessments, translated content, or populations for which evidence does not exist. Geographic region and language/cultural equivalence are separate questions and are not interchangeable.

Future validation requirements include: non-native English speakers; linguistic variation; translated assessments; and cultural response patterns across APAC, the Middle East and continental Europe. Each will be validated in its own right before any claim is made.

17. Regulatory and Responsible-AI Considerations

Athir is designed to inform human decision-making. This positioning does not by itself remove regulatory obligations from Athir or from deploying organisations. Regulatory applicability depends on jurisdiction, intended use, deployment configuration, the client's decision process and the candidate population. Organisations should obtain advice concerning their specific deployment.

European Union. Athir's decision-support design, with human review of outputs, is intended to support responsible use; whether GDPR Article 22 (automated decision-making) applies depends on the deployment configuration and the client's decision process, and Article 22 is only one part of broader GDPR obligations (lawful basis, transparency, data minimisation, purpose limitation, retention, access and rectification, erasure, objection, profiling safeguards, data protection impact assessment, processors and international transfers). Separately, the EU AI Act identifies certain AI systems used in recruitment and candidate evaluation as potentially high-risk depending on intended purpose and deployment; human review alone does not automatically place a system outside that regime. Athir's regulatory classification therefore depends on each deployment, and applicable organisations should undertake a formal EU AI Act classification assessment. Athir does not claim completed compliance with the EU AI Act; a conformity roadmap (risk management, data governance, technical documentation, logging, transparency, instructions for use, human oversight, accuracy, robustness, cybersecurity, quality management, post-market monitoring, incident management, deployer responsibilities and candidate notification) is being developed.

United States and United Kingdom. Athir's fairness methodology is designed with reference to the Uniform Guidelines on Employee Selection Procedures and to the protected characteristics framework of the Equality Act 2010. Reference to these frameworks is methodological; it is not a claim of compliance, which depends on each deployment and is a matter for the deploying organisation and its counsel.

United Arab Emirates. The scoring framework is designed with reference to the equal-treatment principles of Federal Decree Law No. 33 of 2021. Demographic attributes are not explicit scoring inputs; differential performance by untested characteristics has not been empirically established.

18. Governance and Model Versioning

The Athir competency framework, Behavioural Intelligence Engine, scenario bank and scoring rubrics are maintained under documented version control. Every reported statistic is associated with the system version that generated it. Material changes (new model or provider, scoring prompt or rubric changes, competency changes, benchmark or scenario changes, threshold or weighting changes, new languages, new candidate populations) trigger: documentation, regression testing, fairness testing, reliability review, validity-impact review, approval, and version increment. A material change to the production scoring engine triggers an impact assessment to determine whether revalidation is required.

Validation Study (this document): version 1.0, approved for publication, next review in the March 2027 audit cycle.

Bias & Fairness Audit: version 1.0, approved for publication, next review in the March 2027 audit cycle.

Competency framework: maintained in the Athir Governance Register, available to clients on request, and reviewed on material change.

Behavioural Intelligence Engine: maintained in the Athir Governance Register, available to clients on request, and reviewed on material change.

Scenario bank: maintained in the Athir Governance Register, available to clients on request, and reviewed on material change.

Framework, engine and scenario versions are maintained in the Athir Governance Register and provided to clients on request.

19. Ongoing Validation Programme

Athir operates an annual validation cycle covering, for every deployed version: criterion-related validity (when outcome data exists); reliability; score distributions by Target Benchmark and Business Function; subgroup performance; adverse-impact monitoring; model drift; scenario performance and scenario functioning; accessibility; candidate feedback and challenge logs; benchmark calibration; and model/version change impact assessment.

The programme proceeds through defined maturity stages: methodological design; internal technical validation; synthetic fairness testing (complete, preliminary); real-world validation; real-world fairness monitoring; independent external review (intended for the 2027 audit cycle; no reviewer has been appointed, and none will be implied until one is engaged and named); and ongoing post-deployment monitoring. Athir does not describe a stage as complete until its evidence exists.

20. Evidence Maturity Statement

The following states Athir's current evidence position for each evidence area.

Content validity: the framework is documented and formal evidence is in progress. The evidence base is internal design documentation; the next stage is job analysis and independent SME panel review.

Construct and internal structure: not yet established, with nothing published. The next stage is a pre-registered factor-analytic study.

Criterion-related validity: not yet established, with no published coefficient. The next stage is a real-world outcome study (planned Stage 4).

Reliability and measurement precision: partially documented. The evidence base is deterministic score mapping and rubric versioning; the next stage is score-level reliability and engine repeatability studies.

Engine-versus-expert agreement: not yet established, with nothing published. The next stage is an engine-versus-expert agreement study.

Synthetic fairness testing: completed internally and preliminary. The evidence base is 2,847 synthetic responses (see the Bias & Fairness Audit); the next stage is annual retest with unadjusted and adjusted reporting.

Real-world demographic fairness: not yet established, with insufficient real-world data. The next stage runs from the 2027 audit cycle onward, subject to sample sizes.

Independent validation: not yet commissioned, with no evidence published. The next stage is the intended independent review in the 2027 audit cycle.

21. Glossary

Criterion-related validity
The relationship between assessment scores and external performance criteria. It divides into predictive designs (assessment before criterion measurement) and concurrent designs (both measured at approximately the same time). Athir does not currently publish a criterion-related validity coefficient.
Predictive validity
A criterion-related design in which the assessment is administered before the criterion performance is observed. Retrospective or concurrent evidence must not be described as predictive validity.
Reliability
The consistency of scores across repeated administrations, internal indicators, or raters. Reliability at the level scores are interpreted is planned; no coefficient is currently published.
Standard error of measurement
An estimate of the precision of an individual score. Not yet published for Athir; until it is, scores should be interpreted at band level, not as precise points.
Adverse impact
A substantially different selection rate that disadvantages a protected group. Athir does not select candidates, so adverse-impact analysis requires an operationalised selection threshold defined by the deploying organisation.
Four-fifths rule
A screening heuristic (not a legal standard or proof of compliance) under which a group's selection rate below 80% of the highest-selected group's rate signals a need for further review.
Synthetic fairness testing
Preliminary, pre-deployment fairness testing using generated responses with demographic labels. It cannot by itself demonstrate fairness in real human populations.
Composite score
The unweighted mean of the eight competency pillar scores; a documented design decision rather than an empirically derived weighting.
Target Benchmark
The organisational level at which the executive is expected to operate: Enterprise Executive, Business Unit Executive, Functional Executive, or Operational Leader.
Evidence maturity
The stage of validation an evidence area has reached, from methodological design through internal testing, synthetic testing, real-world validation, and independent review.

22. References

External scientific references support Athir's methodology choices. They are not evidence of Athir-specific performance, and Athir never uses external literature to imply that it achieved a statistic. Athir's own validation evidence is proprietary and, where it exists, is available under NDA as part of the Technical Validation Report.

  • American Educational Research Association, American Psychological Association & National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing. Washington, DC: AERA.
  • Society for Industrial and Organizational Psychology (2003). Principles for the Validation and Use of Personnel Selection Procedures (4th ed.). Bowling Green, OH: SIOP.
  • Uniform Guidelines on Employee Selection Procedures (1978). 43 FR 38290; 29 CFR Part 1607.
  • Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334.
  • Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428.
  • Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology. Psychological Bulletin, 124(2), 262–274. Note: cited as methodological background on selection research generally; Athir does not use this source to generate comparative validity claims for its own assessment.
  • Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 (Artificial Intelligence Act).

23. Contact and Technical Documentation

Enterprise clients, procurement teams, legal counsel and qualified technical reviewers may request the Athir Technical Validation Report, which holds the methodological detail behind this document: study designs, sample descriptions, coefficient tables with confidence intervals, reliability tables, subgroup analyses, statistical assumptions, norm development, cut-score documentation, sensitivity analyses and model versions. Where a result does not yet exist, the report marks it as evidence pending rather than presenting an invented figure.

Requests are handled under NDA through the Contact Us page and are reviewed by the Athir Research and Compliance function.

Cookie preferences

Athir uses strictly necessary cookies to operate the platform. Analytics and marketing cookies are used only with your consent. Read the Cookie Policy.