Methodology and Transparency Disclosure — Airline Vertical Pilot
List: The 10 Best Business Class Airlines in Europe for 2026
Methodology version: v1.0 — see /strategy/methodology-v1.md
Category Fit rubric: airline v1.1 — see /strategy/airline-category-fit-rubric-v1.1.md
Run type: Pilot run #3 — first list in the airline vertical, first list against a purpose-built rubric rather than a retrofitted one
Date: 17 August 2026
Editor: SeatAndSuite Editorial
Commercial disclosures: None. SeatAndSuite is pre-revenue; no commercial relationship, sponsorship, affiliate arrangement or pre-publication review for any airline listed.
1. Why this pilot exists
Pilots #1 and #2 tested the methodology across price tiers within a single vertical (London family hotels, budget and luxury). Both produced narrower score distributions than expected — 1.28 and 1.88 points respectively across ten entries — suggesting the manual LCI proxy smooths real signal.
Pilot #3 tests three things the hotel pilots could not:
- Does the five-pillar framework transfer to a second vertical without modification? Methodology v1.0 §8 asserts that it does — "the five pillars are constant, only the inputs change." This is the first test of that claim.
- Does a rubric written before the list perform better than one reverse-engineered after it? Both hotel pilots produced rubric additions retrospectively. This list was scored against a rubric specified first.
- Does the distribution widen in a vertical with more objectively measurable product differences? Hypothesis going in: yes. Aircraft cabins are physically specified, publicly documented and vary far more than hotel family programmes.
It also tests a structural case hotels never posed: a single entity operating two materially different products under one fare-class name. The reader-facing decision — score long-haul and intra-Europe business class together — was set by the commissioning brief; the methodological consequences are documented in full below.
2. Scope and inclusion criteria
In scope. Airlines headquartered in Europe (including Turkey and the United Kingdom) operating a cabin sold as business class, on long-haul routes, intra-Europe routes, or both. Leisure and hybrid carriers are explicitly included per the commissioning brief.
Out of scope. Non-European carriers operating European routes — Qatar Airways, Emirates, Singapore Airlines and others were excluded by definition, not by score. Several would rank above every airline here.
Excluded for absence of product. Norse Atlantic Airways and PLAY operate no business class cabin. Norse's Premium is a 43-inch-pitch recliner premium economy product; scoring it against lie-flat business class would compare different products. Icelandair's Saga Premium is likewise a recliner and excluded on the same basis. This exclusion is worth noting because the commissioning brief explicitly asked for leisure and hybrid carriers to be considered — they were considered and failed the product threshold, rather than being overlooked. La Compagnie, the only European all-business-class carrier, was assessed and scored (16th).
Assessed and scored but outside the top ten: ITA Airways, Aer Lingus, SAS, Austrian Airlines, LOT Polish Airlines, La Compagnie, TAP Air Portugal. Full scores in §5.
Assessed and not scored for insufficient data: Air Europa, Brussels Airlines, Croatia Airlines, TAROM, Air Serbia. None were plausible top-ten candidates and the data volume did not justify the scoring effort. This is a resource decision, disclosed rather than concealed.
3. Pilot proxies (consistent with pilots #1 and #2)
| Pillar | Weight | Full v1.0 spec | Pilot proxy used here |
|---|---|---|---|
| LLM Citation Index | 35% | 50-prompt × 4-model weekly run, ~200 runs per category | Recurrence and positioning across the answer-engine sources most commonly cited by ChatGPT, Claude, Perplexity and Google AI Overviews for European business class queries, manually scored 0–10 with sentiment adjustment |
| Aggregated Review Sentiment | 25% | Multi-source, recency- and depth-weighted, NLP-scored | Editorial reading of Skytrax, AirlineRatings and TripAdvisor airline sections plus long-form review sites, weighted toward business-cabin-specific content |
| Star Ratings & Awards | 15% | OTA volume-weighted average + official rating + recent awards | Skytrax World Airline Awards 2025 placement and category wins; AirlineRatings World's Best Airlines 2026 placement; Skytrax star rating; APEX standing |
| Search Visibility & Authority | 15% | Keyword position bank, snippet/KP, DR-weighted backlinks, Wikipedia | Manual SERP assessment against a European-business-class keyword set plus brand-domain authority signal |
| Category Fit | 10% | Editorial 0–10 score against published rubric | Same — fully aligned with v1.0 spec, scored against airline rubric v1.1 |
As in both hotel pilots, Category Fit is the only pillar fully aligned with the v1.0 specification. Rankings within ±2 positions should be treated as effectively tied under the pilot proxies. The gap between SWISS (8.22) and Virgin Atlantic (8.20) is 0.02 points and is not a real distinction.
3.1 What is different about the airline proxies
Three proxy behaviours differ from the hotel pilots and should be understood before the numbers are read.
Search Visibility is more punishing in this vertical. Airline brand domains carry vastly more authority than hotel domains — British Airways scores 9.4 and Aegean 6.2 on the same scale where the widest hotel gap in pilot #2 was 9.5 to 6.5. Because Search Visibility carries 15%, this systematically advantages flag carriers of large economies over excellent airlines from small ones. This is a real effect in AI-search citation behaviour, not a proxy artefact — but it means the ranking partly measures national economic scale.
Awards data is unusually current and unusually stale at the same time. AirlineRatings' 2026 edition is current. Skytrax's is not: the World Airline Awards run an annual September cycle and the 2025 edition remains the current one until 18 September 2026. Every airline's Awards pillar score here is built on a mix of 2025 and 2026 data, and will re-set within five weeks of publication. This is disclosed on the list itself and is the reason the next refresh is scheduled for 21 September rather than a month out.
Review sentiment is more reliable here than in hotels. Airline review volume per entity is an order of magnitude higher than hotel review volume, and Skytrax and AirlineRatings both publish structured cabin-specific assessments. The Sentiment pillar is the pillar we trust most in this list, which is the reverse of the hotel pilots.
4. Category Fit — how the two-segment score was constructed
Full rubric in airline-category-fit-rubric-v1.1.md. Summary of the mechanics as applied:
Category Fit = (Long-haul sub-score × 0.6) + (Intra-Europe sub-score × 0.4)
Where a carrier does not operate a segment, that sub-score is recorded N/A and the remaining sub-score carries 100%. Two airlines in the scored set are affected: Virgin Atlantic (no intra-Europe business class) and La Compagnie (no intra-Europe operation at all). Both are flagged on the published list.
This handling gives long-haul-only carriers a structural advantage — they cannot be dragged down by a weak short-haul cabin because they do not operate one. Virgin Atlantic's Category Fit of 8.8 is the highest on the list and is not like-for-like with Turkish Airlines' 8.3, which is a blend of a 7.5 long-haul and a 9.5 short-haul.
Sensitivity analysis on Virgin Atlantic:
| Treatment of absent segment | Category Fit | Weighted total | Position |
|---|---|---|---|
| N/A, long-haul carries 100% (used) | 8.80 | 8.20 | 4th |
| Typical European short-haul product (6.0) | 7.68 | 8.09 | 4th |
| Scored as zero | 5.28 | 7.85 | 6th |
The N/A treatment is worth 0.11 points against a like-for-like comparator — real, but not position-changing. Scoring the absent product as zero would move Virgin to sixth, and would be the wrong answer: it penalises a network decision rather than a product failure. We consider the N/A treatment correct and the disclosure mandatory.
Fleet-exposure weighting was applied to every long-haul sub-score. The hard-product component is (best-seat score × fleet share) + (legacy-seat score × legacy share) rather than the best seat in the fleet:
| Airline | Flagship seat | Approx. long-haul fleet exposure | Effect on long-haul sub-score |
|---|---|---|---|
| Lufthansa | Allegris | 19 aircraft; A350 retrofit begins 2027 | 8.9 unweighted → 7.0 weighted |
| Turkish Airlines | Crystal | 0% — not in commercial service | 8.7 unweighted → 7.5 weighted |
| SWISS | Senses | A350s only; A330 retrofit under way | 9.0 unweighted → 8.2 weighted |
| Air France | Safran Versa suite | ~13 of 40 Boeing 777-300ERs | 8.9 unweighted → 8.5 weighted |
| British Airways | Club Suite | ~60–70% of long-haul fleet | 8.6 unweighted → 7.8 weighted |
And here is the uncomfortable result: re-running the full model with fleet-exposure weighting switched off changes nobody's position. The top eleven come out in exactly the same order — Turkish 8.93, Air France 8.55, SWISS 8.27, Virgin 8.20, British Airways 8.03, Lufthansa 7.94, Iberia 7.45, KLM 7.31, Finnair 7.08, Aegean 6.76, ITA 6.62. The largest single movement is Lufthansa gaining 0.11 points.
This was the rubric's most considered rule and the one we expected to do the most work. It is directionally right and numerically inert, for the same reason set out in §6: Category Fit carries 10%, so even a rule that moves a long-haul sub-score by 1.9 points moves the total by 0.11. We are publishing this rather than quietly dropping the claim, because it is the single strongest piece of evidence for the v1.2 weighting recommendation.
Seat-selection surcharge deduction was applied to Lufthansa and SWISS (−1.5 and −1.0 respectively on the seat criterion within the long-haul sub-score). This is a deliberate rubric position, not a neutral measurement, and is stated as such in the rubric §3.1. Its arithmetic effect is small for the same reason as above: +0.027 for Lufthansa and +0.018 for SWISS if reversed, changing no positions. The surcharge's real effect on these airlines' rankings runs through the Sentiment and LCI pillars, where it is the dominant negative theme in coverage of Allegris — which is a 60%-of-weight effect, not a 10% one.
5. Per-airline scoring
Scores 0–10 per pillar. Weighted total uses methodology v1.0 weights (35/25/15/15/10). Category Fit shown decomposed into its two sub-scores.
| # | Airline | LCI (35%) | Sentiment (25%) | Awards (15%) | Search (15%) | CF-LH | CF-SH | Cat Fit (10%) | Weighted | Confidence |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Turkish Airlines | 9.2 | 8.3 | 9.5 | 8.7 | 7.5 | 9.5 | 8.30 | 8.86 | 4 |
| 2 | Air France | 8.8 | 8.4 | 8.4 | 8.8 | 8.5 | 6.3 | 7.62 | 8.52 | 5 |
| 3 | SWISS | 8.4 | 8.5 | 8.0 | 8.0 | 8.2 | 6.5 | 7.52 | 8.22 | 3 |
| 4 | Virgin Atlantic | 7.9 | 8.5 | 8.0 | 8.2 | 8.8 | N/A | 8.80 | 8.20 | 5 |
| 5 | British Airways | 8.2 | 7.4 | 7.6 | 9.4 | 7.8 | 6.0 | 7.08 | 7.98 | 4 |
| 6 | Lufthansa | 8.0 | 7.6 | 7.5 | 9.0 | 7.0 | 5.8 | 6.52 | 7.83 | 3 |
| 7 | Iberia | 7.2 | 7.8 | 7.2 | 7.8 | 7.8 | 6.5 | 7.28 | 7.45 | 4 |
| 8 | KLM | 6.6 | 7.9 | 7.3 | 8.4 | 7.0 | 6.2 | 6.68 | 7.31 | 3 |
| 9 | Finnair | 6.9 | 7.7 | 6.5 | 7.2 | 6.8 | 7.0 | 6.88 | 7.08 | 4 |
| 10 | Aegean Airlines | 5.8 | 8.7 | 7.0 | 6.2 | 4.0 | 8.5 | 5.80 | 6.76 | 4 |
| 11 | ITA Airways | 6.2 | 7.5 | 5.8 | 6.6 | 7.9 | 6.0 | 7.14 | 6.62 | 3 |
| 12 | Aer Lingus | 5.6 | 7.2 | 5.4 | 7.0 | 7.0 | 6.8 | 6.92 | 6.31 | 4 |
| 13 | SAS | 5.5 | 7.4 | 5.7 | 6.8 | 7.0 | 5.8 | 6.52 | 6.30 | 3 |
| 14 | Austrian Airlines | 5.4 | 7.3 | 5.7 | 6.4 | 5.5 | 6.8 | 6.02 | 6.13 | 4 |
| 15 | LOT Polish Airlines | 5.0 | 7.0 | 6.3 | 6.0 | 6.8 | 5.5 | 6.28 | 5.97 | 4 |
| 16 | La Compagnie | 5.5 | 8.0 | 4.5 | 4.2 | 6.5 | N/A | 6.50 | 5.88 | 3 |
| 17 | TAP Air Portugal | 5.2 | 6.6 | 5.0 | 6.6 | 6.8 | 5.5 | 6.28 | 5.84 | 4 |
Score distribution across the published top ten: 8.86 − 6.76 = 2.09 points. Score distribution across the full scored set of seventeen: 8.86 − 5.84 = 3.02 points.
6. Segment-split rankings
Published on the list itself, reproduced here with the calculation basis.
Long-haul only — Category Fit = long-haul sub-score, intra-Europe discarded:
| # | Airline | Score | vs combined |
|---|---|---|---|
| 1 | Turkish Airlines | 8.78 | — |
| 2 | Air France | 8.61 | — |
| 3 | SWISS | 8.29 | — |
| 4 | Virgin Atlantic | 8.20 | — |
| 5 | British Airways | 8.05 | — |
| 6 | Lufthansa | 7.88 | — |
| 7 | Iberia | 7.50 | — |
| 8 | KLM | 7.34 | — |
| 9 | Finnair | 7.08 | — |
| 10 | ITA Airways | 6.70 | +1 (in) |
| 11 | Aegean Airlines | 6.58 | −1 (out) |
Intra-Europe only — Category Fit = intra-Europe sub-score, long-haul discarded. Virgin Atlantic and La Compagnie drop out entirely:
| # | Airline | Score |
|---|---|---|
| 1 | Turkish Airlines | 8.97 |
| 2 | Air France | 8.39 |
| 3 | SWISS | 8.12 |
| 4 | British Airways | 7.87 |
| 5 | Lufthansa | 7.75 |
| 6 | Iberia | 7.37 |
| 7 | KLM | 7.26 |
| 8 | Finnair | 7.09 |
| 9 | Aegean Airlines | 7.04 |
| 10 | ITA Airways | 6.50 |
The most important finding in this pilot is in that third table. Scoring Category Fit exclusively on the intra-Europe product moves Aegean — which has, by broad reviewer consensus, the second-best short-haul business class in Europe — from tenth only to ninth. Lufthansa, whose intra-Europe product is a blocked middle seat with thin catering, still ranks fifth.
This is not an editorial disagreement with the algorithm. It is the algorithm working exactly as specified and revealing that at a 10% weight, Category Fit cannot answer a product question. Sixty-five percent of the weight sits in citation volume, search authority and awards standing — three signals that all reward brand scale. For a category question of the form "which brand is best known and best regarded", that is correct. For one of the form "which product is better", it is not.
7. Pilot #1 vs #2 vs #3 — findings
7.1 Score-distribution comparison
| Pilot | Category | Entries | Spread (top 10) | Spread (full scored set) |
|---|---|---|---|---|
| #1 | London budget family hotels | 10 | 1.28 | not published |
| #2 | London luxury family hotels | 10 | 1.88 | not published |
| #3 | European business class airlines | 10 (17 scored) | 2.09 | 3.02 |
The hypothesis held. Pilot #2 predicted that richer third-party coverage and more differentiating rubric criteria widen the distribution. Airlines have both, in greater degree than luxury hotels, and produced the widest top-ten spread yet — a 63% improvement on pilot #1 and 11% on pilot #2.
The target set in pilot #2 was a 3.0–4.0 point spread under the full automated stack. The full seventeen-airline scored set already reaches 3.02 under the manual proxy, which suggests the proxy smoothing is less severe in this vertical than in hotels — most likely because cabin products are physically specified and publicly documented in a way that hotel family programmes are not.
Recommendation carried forward: publish the full scored set spread alongside the top-ten spread in all future lists. Pilots #1 and #2 published only the top ten, which understates the discrimination the methodology actually achieves and makes the pilots look worse than they were. Retrofit this figure into the pilot #1 and #2 methodology documents at their next refresh.
7.2 Rubric-first versus rubric-retrofit
Pilots #1 and #2 both produced rubric criteria after scoring — three budget criteria and six luxury criteria respectively, all logged as v1.1 recommendations. Pilot #3 wrote the rubric first.
The rubric-first approach produced two criteria that would never have emerged retrospectively: fleet-exposure weighting, and the fare-class integrity criterion for intra-Europe. Both required thinking about the category's failure modes in the abstract rather than reacting to the candidate set. Neither moved a single position under v1.0 weights (§4) — which is a finding about the weights, not about the criteria.
It also produced one criterion that did not earn its place: on-time performance, specified in the rubric's consistency component and in methodology v1.0 §8. OTP data across the scored set clustered tightly enough that it discriminated almost nothing, and the publicly available data is inconsistently defined between carriers and regulators. Recommendation: demote OTP to a tie-breaker rather than a scored criterion, or drop it.
Recommendation for v1.2: write the rubric before the list, always. Formalise this as a step in the list-production process — no list is scored until its category rubric is written and reviewed.
7.3 Where the proxy disagreed most with the rubric
Consistent with the Athenaeum anomaly in pilot #2, one entity shows a large gap between its Category Fit score and its ranking:
ITA Airways. Long-haul Category Fit 7.9 — the third-highest on the entire list, above Turkish Airlines, British Airways, Lufthansa and Finnair. Ranks eleventh overall on a total of 6.62. The gap is driven entirely by LCI (6.2), Awards (5.8) and Search (6.6), all of which are legacy effects of the Alitalia collapse and brand reset rather than assessments of the current product.
This is the canary for the airline vertical, exactly as the Athenaeum is for luxury hotels. Logged prediction: ITA moves up materially when the full LCI harness goes live and product-quality prompt clusters are tracked separately from brand-recognition clusters. If it does not move, the harness is mis-tuned for the airline vertical.
Secondary anomaly — Aer Lingus. AerSpace on the A321LR and XLR is a genuinely lie-flat intra-Europe business class, one of only three such products in Europe. It is almost entirely absent from answer-engine outputs for European business class queries. LCI proxy 5.6 against an intra-Europe sub-score of 6.8. A clear content-gap opportunity and a strong B2B Pulse pitch.
7.4 Confidence Score behaviour
Average confidence across the published top ten: 3.9, against 4.4 in pilot #2 and a higher figure in pilot #1. This is the lowest-confidence list SeatAndSuite has published, and the reason is structural rather than a data problem: four of the top ten are in active product transition.
| Airline | Confidence | Reason for cap |
|---|---|---|
| SWISS | 3 | Flagship cabin entered service within 12 months; fleet exposure below 30% |
| Lufthansa | 3 | Fleet exposure below 30%; A380 retrofit began April 2026 |
| KLM | 3 | Certification event has made near-term product allocation unpredictable |
| Turkish Airlines | 4 | Flagship product announced but not in service; wide variance in seat assigned |
| British Airways | 4 | Bifurcated fleet; experience unpredictable at individual-booking level |
The airline vertical is structurally lower-confidence than hotels and readers should be told so. A hotel's product changes on a renovation cycle measured in years. An airline's changes with every aircraft assignment. The 12-month product-transition cap (rubric §2.3) fired on two airlines at first application, which suggests it is calibrated about right — a rule that never fires is not a rule.
Recommendation: add a list-level "product stability" indicator to airline lists, distinct from per-entry confidence, telling readers how much of the ranking is likely to move within two refreshes. This list would score low on it.
7.5 Criteria found insufficient
On-time performance — see §7.2. Insufficient discrimination, inconsistent public definitions.
Lounge assessment needs decomposition. Ground experience carries 20% of the long-haul sub-score, and lounge quality dominates it, but "lounge quality" as a single 0–10 judgement conflates home-hub flagship quality, outstation coverage, crowding and food. Virgin Atlantic's Clubhouse and Turkish's Istanbul lounge score similarly on a single scale while being good at completely different things. Recommendation: split into home-hub quality, outstation coverage, and crowding/capacity.
Wi-Fi needs to be scored on fleet coverage, not availability. Same failure mode as the seat: "has Wi-Fi" is a fleet-share question, not a binary. The rubric already specifies coverage percentage; the pilot scored it loosely because reliable per-fleet coverage data was not available for most carriers. Flag as a data-sourcing gap.
8. Editorial decisions log (anchor #4)
Per methodology v1.0 §4, the editor cannot alter algorithmic scores. Decisions taken in this pilot:
Aegean Airlines retained at #10 despite having no lie-flat long-haul product. The algorithm placed it tenth on combined scoring. Editorial view is that this is defensible under the commissioned scope (both segments scored together) but potentially misleading to a reader shopping long-haul. Resolution: retained at the algorithmic position, with the caveat stated explicitly in the entry itself and the long-haul-only reordering published directly beneath the list. No score was altered. This is the correct application of anchor #4 — the editor's tool is disclosure, not adjustment.
Norse Atlantic, PLAY and Icelandair excluded on product-threshold grounds. All three were assessed; none operates a lie-flat business class cabin. Documented in §2 because the commissioning brief specifically requested leisure and hybrid carriers be included, and the reader is entitled to know they were considered.
La Compagnie scored rather than excluded. The only European all-business-class carrier. Ranks sixteenth largely because Awards (4.5) and Search Visibility (4.2) — 30% of the weight combined — are structurally unavailable to an airline of its size. Editorial note published on the list: as a value proposition on the routes it flies, it beats several airlines ranked above it. This is a known limitation of a methodology that weights brand scale, disclosed rather than corrected.
Turkish Airlines' Crystal suite scored at zero fleet exposure. It has been announced, photographed and covered extensively, and was not in commercial service as of publication. Under fleet-exposure weighting it contributes nothing to the long-haul sub-score. Announced products do not score.
Air Europa, Brussels Airlines, Croatia Airlines, TAROM and Air Serbia assessed but not scored. None plausible for the top ten; data volume did not justify the effort. Disclosed as a resource decision.
No airline was contacted before publication. Consistent with pilots #1 and #2. Dispute-resolution process remains an open item from methodology v1.0 §11 and is now more urgent — airlines have larger communications functions than individual hotels and are more likely to contest a rank.
9. B2B Pulse audit angle (anchor #3) — airline edition
The airline vertical is a materially better Pulse market than hotels, for three reasons: airline marketing budgets are larger, the buying centre is more concentrated (one CMO desk covers the whole network rather than one per property), and airlines already buy competitive-intelligence products, so the category does not need to be explained.
Specific audit hooks surfaced by this pilot:
ITA Airways. The clearest pitch on the list. Third-best long-haul product, eleventh-place citation and search visibility. The audit writes itself: "Your product scores in the top three in Europe. AI search engines rank you eleventh. Here is the prompt-level gap and here is what is causing it." Expected finding: LLMs are citing legacy Alitalia-era sources and generic aggregator content rather than current ITA product pages.
Aer Lingus. Operates one of only three genuinely lie-flat intra-Europe business class products in Europe and is close to invisible in answer-engine outputs for the query it should own. A content and schema gap, not a product gap. High-conviction quick win.
Aegean Airlines. Highest review sentiment of any European carrier (8.7) and the second-best intra-Europe product, against a 5.8 LCI proxy and 6.2 search visibility. Aegean is under-cited relative to how much its customers like it. The audit angle is sentiment-to-citation conversion.
Lufthansa. The most commercially interesting audit and the hardest sell. Allegris generates enormous positive coverage; the seat-selection surcharge generates enormous negative sentiment; the two are entangled in every answer-engine output about the product. A prompt-level sentiment analysis showing exactly where the surcharge narrative enters the citation graph would be genuinely valuable to their team, and genuinely unwelcome.
KLM. Time-sensitive. The A350 certification story is currently entering the citation graph and will shape how LLMs describe KLM business class for months after it is resolved. A before/during/after citation time series is exactly the Pulse "trend over time" view and is a live demonstration of why weekly tracking matters. Recommendation: begin tracking KLM A350 citation sentiment now, before the pitch, so the time series exists when the conversation happens.
Turkish Airlines. Crystal launch is imminent. Pre-launch, launch and post-launch citation tracking is the single best demonstration case available in European aviation for the next twelve months. Same recommendation: start the time series before the pitch.
10. Limitations and what could go wrong
The proxy is not the harness. Three of five pillars are manual proxies. LCI in particular is an editorial reading of what answer engines appear to cite, not a measured 200-run dataset. Positions within ±2 are not meaningfully distinguished.
Fleet exposure figures are point-in-time and were changing during production. Retrofit programmes at Air France, SWISS, British Airways and Lufthansa are all active. Every fleet-share figure in this list is approximate, sourced from airline disclosures cross-checked against enthusiast fleet trackers, and will be wrong within a quarter. This is the fastest-decaying data in the list.
The Awards pillar is mid-cycle. Skytrax 2025 data will be superseded on 18 September 2026. Any airline whose 2026 result diverges from 2025 will move. Turkish Airlines' Best Airline in Europe title is load-bearing for its first-place finish.
Brand scale contaminates a product ranking. Documented at length in §6. Fifty percent of the weight — LCI at 35% and Search Visibility at 15% — rewards being well known. For a list titled "best business class airlines", that is a partial answer to a different question. The v1.2 weighting recommendation addresses it; until then, readers should treat the segment-split tables and the Category Fit column as the product-quality view.
Two entries have an N/A sub-score. Virgin Atlantic and La Compagnie are not scored on a like-for-like basis with the rest. The sensitivity analysis in §4 quantifies the effect for Virgin (would fall to fifth with a typical European short-haul product). Disclosed, quantified, not corrected.
We have not flown these products in the review period. Consistent with methodology v1.0 §10 — SeatAndSuite synthesises online signals and does not conduct site visits or review flights. The Category Fit scores are built from published seat specifications, aeroLOPA seat maps, airline disclosures and review aggregation, not from first-hand assessment.
The seat-selection-surcharge deduction is a value judgement. It is defensible and it is disclosed, but it is not a neutral measurement. An airline could reasonably argue that unbundling seat selection is a fare-structure decision, not a product defect. We disagree, publicly, and have said why. Readers who disagree with us can add 0.027 to Lufthansa's total and 0.018 to SWISS's; neither changes a position.
11. Recommendations from this pilot
Carried into the next methodology review, in priority order:
-
Upweight Category Fit to 20% for product-led categories, with LCI reduced to 30%. The strongest finding in this pilot, supported by three independent pieces of evidence: fleet-exposure weighting moves no positions (§4), the surcharge deduction moves no positions (§4), and scoring Category Fit exclusively on intra-Europe moves Aegean by one place (§6). Every product-quality mechanism in the rubric is currently numerically inert.
Re-running the model at 30/25/15/15/20 produces: Turkish 9.22, Air France 8.84, Virgin Atlantic 8.69 (up to 3rd), SWISS 8.55, British Airways 8.28, Lufthansa 8.08, Iberia 7.82, KLM 7.65, Finnair 7.43, Aegean 7.05, ITA Airways 7.02 (from 0.14 behind to 0.03 behind). The reweighting moves the airlines with the best products up and compresses the gap between product quality and brand scale — which is what the list title promises. Would be the first weight change since v1.0 and must go through the public changelog. Test on the next two lists before adopting.
-
Write the category rubric before the list, always. Formalise as a production gate. Rubric-first produced two criteria in this pilot that could not have emerged retrospectively (§7.2).
-
Publish full-scored-set distribution alongside top-ten distribution in every list, and retrofit the figure into pilots #1 and #2.
-
Demote on-time performance to a tie-breaker or drop it. Insufficient discrimination, inconsistent public definitions.
-
Decompose the lounge criterion into home-hub quality, outstation coverage, and crowding/capacity.
-
Add a list-level product-stability indicator to airline lists, distinct from per-entry confidence.
-
Source per-fleet Wi-Fi coverage data. Currently a known gap; the rubric specifies coverage percentage and the pilot could not supply it.
-
Build the dispute-resolution process before the next airline list. Open since methodology v1.0 §11 and now urgent — airlines contest rankings more readily than hotels.
-
Begin LCI time-series tracking on KLM and Turkish Airlines now, ahead of the A350 certification resolution and the Crystal launch respectively, so the time series exists when the Pulse conversation happens.
12. Changelog
| Version | Date | Change |
|---|---|---|
| 1.0 | 17 Aug 2026 | Initial publication. Pilot run #3, first airline-vertical list. First application of airline Category Fit rubric v1.1, fleet-exposure weighting, two-segment Category Fit split and the 12-month product-transition confidence cap. |
Next refresh: 21 September 2026, immediately following the 2026 Skytrax World Airline Awards ceremony on 18 September at The Langham, London.