Lunartulip Lab · AlphaMap Research Note #001

The Projection Is Part of the Signal

From AI-industry semantics to cross-sectional stock rankings

18 September 2026 · Exploratory methods study

The choice of industry labels can determine much of a graph-based stock score before any returns are examined.

We built an end-to-end graph-to-ranking pilot and found a precise measurement effect: in this coarse role projection, weighted degree is exactly an exposure count, while 31 of 39 Katz scores admit a group-size explanation. The result provides a practical baseline for testing what richer semantic relationships—and eventually changing industry states—add to a stock signal.

39global graph issuers
26priced stocks
1historical graph vintage

A working path from semantics to a stock ranking

AlphaMap’s first graph experiment converts a maintained AI-industry ontology into company features, mechanical stock rankings and hypothetical portfolio outcomes. We call this repeatable transformation the Graph Signal Factory. Its first output is an interpretable measurement result: a graph score can reproduce assumptions already embedded in the ontology.

01 · INPUTDated semantic memberships
02 · GRAPHExplicit company projection
03 · FEATURECentrality and exposure
04 · SIGNALControlled score and rank
05 · TESTWeights, outcomes and exposures

The input is a historically maintained set of issuer–role and issuer–theme memberships, projected deterministically. It contains 39 global issuers; 26 US-listed names have the archived controls required for the common priced sample. We compute the graph on all 39 before restricting the portfolio universe. The role projection has nine broad labels; a separate theme projection uses seven research-view memberships.

Sharing a label creates a peer relationship in this experiment. It does not establish a supplier contract, a customer dependency or the sign of a cash-flow transmission. The topology makes that distinction visible.

Figure 1. Actual frozen role projection, not a supplier-contract network. Eight components include two isolates. Block density reflects shared role labels; the side panel uses the same issuer order. PNG · PDF

Centrality can have an exposure-only explanation

Thirty-seven of the 39 issuers have only one recorded role. The resulting role graph has eight components. Seven are regular cliques, including two singletons; together they contain 31 issuers. Only the eight-node compute/networking component contains overlapping roles.

For a binary company–role matrix B, let fᵣ be the number of companies assigned to role r and let its weight be qᵣ = 1 + log((N + 1)/(fᵣ + 1)). The projected graph is W = B diag(q) Bᵀ, with its diagonal removed. Its weighted degree is exactly:

strengthᵢ = Σᵣ Bᵢᵣ qᵣ (fᵣ − 1)

This identity holds for all 39 nodes. Weighted degree adds no information beyond weighted role exposures under this construction. Within an isolated homogeneous role group of size m, every node has the same weighted degree d = (m − 1)[1 + log((N + 1)/(m + 1))]. Our excess-over-baseline Katz convention then gives:

k = (I − αW)⁻¹1 − 1
kᵢ = αd / (1 − αd),   α = 0.85 / ρ(W)

The formula reproduces the observed Katz scores of those 31 issuers to numerical precision. The eight-node overlap component is outside this group-size result. These are established linear-algebra identities applied to the actual dataset, rather than a new centrality theorem. Katz (1953).

Figure 2. Two exact representation checks. Weighted degree equals summed weighted role frequencies for all 39 nodes. Katz follows a group-size formula for 31 nodes in regular components; the eight-node compute/networking bridge is explicitly outside that formula. PNG · PDF

Eigenvector centrality is even more concentrated here: all positive mass lies on the nine-node demand/cloud component; each member receives 1/3 under unit-length normalization. The other 30 nodes receive zero. This is a consequence of disconnected components and their spectral radii.

TSMC and Snowflake each have a unique recorded role in this projection and therefore no shared peer. Their zero role scores say nothing about their economic importance. A usable signal factory must distinguish no recorded peer from an economically meaningful low score before converting the lower tail into a short book.

Measure what remains after simple exposures

On the 26-stock priced sample, we compare each raw score with its residual from a joint regression on log market capitalization, research-theme indicators and a theme-membership-count coverage proxy. The design has nine independent columns, including the intercept, leaving a 17-dimensional residual subspace.

Figure 3. In-sample score variance, not explained investment returns. Controls are an intercept, archived log market cap, research-theme indicators and theme-membership count: 9 independent columns, 26 observations, 17 residual degrees of freedom. Zeroes with algebraic causes are identified explicitly. PNG · PDF

The controls span 90.64% of role-Katz and 96.84% of theme-Katz score variance in this sample. These are descriptive, in-sample relationships between features and controls. They do not measure explained investment returns, out-of-sample predictability or a causal contribution from coverage.

Two zero residuals have exact reasons. Theme count is itself a control; role eigenvector is a linear combination of two theme indicators. Their abstention is a correct handling of a design identity. It is not an independent market rejection of either feature.

The practical question for richer graphs is therefore precise: what additional information is created by specific technologies, products or verified economic relations after these simpler exposures are retained as baselines?

Return comparisons test a specification, not an explanation

The frozen scoring rules generate fractional-tie top/bottom 20% allocations. We retain all six features and both raw and residual-score treatments. The illustration below uses the September 9, 2026 open to the September 15 close: five trading sessions on a common 26-stock sample.

Figure 4. September 9 open to September 15 close, 2026; common sample of 26 stocks. Scores are transformed before sorting, so the differences are not causal return decompositions. † Raw role-count ties leave only $1/3 invested per leg; other evaluated legs deploy $1 each. Before costs; no annualization or alpha inference. PNG · PDF

The +8.50-point raw theme-Katz spread becomes +0.63 points under the residual-score sorting rule. Both the stocks and weights change. This is specification sensitivity; the difference cannot be causally assigned to “alpha removed by controls.” The raw spread combines a −2.26-point long-leg contribution with a −10.77-point return on the underlying short basket. The positive spread is not a rising long book.

Score neutralization is also distinct from portfolio neutralization. The nonlinear tail-selection step can reintroduce exposure. The residual role-Katz portfolio retains a one-third long allocation to Memory and a one-third short allocation to Optical/Networking, as well as market beta of +0.257.

Figure 5. Exposures of portfolios formed from residualized scores. Role Katz retains substantial theme imbalance. Theme Katz happens to have zero net theme weights here, but retains size, coverage-proxy and market-beta exposure. Research themes are not GICS sectors. PNG · PDF

The returns are exploratory diagnostics from one retrospectively designed vintage. The preallocated budget is $1 per leg; overlapping ties cancel and the corresponding allocation remains undeployed. Raw role count deploys only $1/3 per leg, so its raw/residual pair is not an equal-deployment comparison. All other evaluated portfolios deploy $1 per leg. A spread is PnL per $1 reference allocation, not a funded portfolio return.

Turn ontology resolution into a falsifiable research programme

The pilot establishes a repeatable path from semantic memberships to auditable stock rankings. Its substantive contribution is to expose where the representation determines the score. This makes the next investment in graph quality testable.

Next studyWhat changesWhat would support incremental value
Projection and resolutionAdd specific products/technologies and typed, evidenced company relationships; track missingness explicitly.Useful variation beyond role counts, group size, coverage and the same eligible universe; comparison with an appropriate structure-preserving null.
Dynamic industry statesAttach comparable, dated state revisions and explicit transmission rules to economic relations.Propagated-state features improve on company-only state and centrality on identical information sets.
Repeated-vintage validationFreeze processing and portfolio rules before new outcomes; include calendars, security lifecycle, risk constraints and costs.Stable performance across independent dates and event clusters, with realistic implementation and predeclared tests.

LLM-assisted extraction is a promising way to populate more detailed semantic inputs. This experiment does not measure its extraction accuracy or investment contribution. Likewise, dynamic state propagation remains a separate, untested return hypothesis. The next series entries will make those contributions explicit rather than infer them from a functioning ranking pipeline.

Study scope and numerical appendix

Design. Retrospective exploratory construction study. The graph records were archived on September 8, 2026; the capitalization/theme controls were archived before the hypothetical September 9 entry. The experiment was designed after its one- and five-session outcomes. Historical availability of inputs does not make the study prospective. Seven foreign listings remain in the graph but are outside this initial US-calendar portfolio test; six US-listed names lack the required archived controls. The sample is curated, not a broad-market universe.

Prices and implementation. Historical Yahoo Finance adjusted prices, retrieved September 17, 2026. Return proxy: adjusted exit close divided by entry open scaled by the entry-day adjustment ratio, minus one. Completed sessions only, fixed cutoff September 16. Market beta uses up to 252 daily returns before entry, with at least 120; the static hedge diagnostic subtracts portfolio beta times the matching SPY holding return. Cost scenarios use 10/25/50bp per absolute dollar of adjusted-value trading notional, not a full corporate-action, financing or stock-borrow ledger.

Inference. Six feature definitions, two score treatments and two mature horizons produce 24 reporting cells: 20 evaluations and four abstentions. They share one entry date. Unrestricted issuer-label permutations are retained as internal diagnostics only; no topology-specific significance is claimed. The planned 20-session follow-up is incomplete at this cutoff and, when realized, will remain part of this selected historical cohort. No annualized alpha, Sharpe ratio, discovery p-value or dynamic-propagation performance is inferred.

All one- and five-session outcomes, with deployed notionals
SessionsFeatureScoringLong / short $Gross PnL, ppStatic beta-hedged, pp25bp cost proxy, pp
1Role strengthRaw1.000 / 1.000-0.83-0.78-1.83
1Role strengthResidual score1.000 / 1.000+0.54+0.53-0.47
1Role KatzRaw1.000 / 1.000-1.88-1.76-2.88
1Role KatzResidual score1.000 / 1.000+0.61+0.66-0.40
1Role eigenvectorRaw1.000 / 1.000-2.63-2.64-3.62
1Role eigenvectorResidual scoreAbstainAbstainAbstain
1Theme KatzRaw1.000 / 1.000+1.26+1.04+0.27
1Theme KatzResidual score1.000 / 1.000+0.96+0.94-0.03
1Role countRaw0.333 / 0.333+0.91+0.92+0.58
1Role countResidual score1.000 / 1.000-0.67-0.70-1.67
1Theme countRaw1.000 / 1.000-0.90-1.08-1.90
1Theme countResidual scoreAbstainAbstainAbstain
5Role strengthRaw1.000 / 1.000-1.36-1.16-2.33
5Role strengthResidual score1.000 / 1.000-2.61-2.64-3.57
5Role KatzRaw1.000 / 1.000-0.73-0.24-1.69
5Role KatzResidual score1.000 / 1.000-2.31-2.08-3.26
5Role eigenvectorRaw1.000 / 1.000-0.40-0.47-1.36
5Role eigenvectorResidual scoreAbstainAbstainAbstain
5Theme KatzRaw1.000 / 1.000+8.50+7.62+7.54
5Theme KatzResidual score1.000 / 1.000+0.63+0.55-0.32
5Role countRaw0.333 / 0.333+1.33+1.36+1.01
5Role countResidual score1.000 / 1.000-3.01-3.16-3.97
5Theme countRaw1.000 / 1.000+3.61+2.91+2.65
5Theme countResidual scoreAbstainAbstainAbstain

Amounts are per $1 reference. Beta-hedged and cost-stressed figures are separate diagnostics, not a jointly financed net strategy. Abstention preserves zero exposure rather than forcing a rank.

Intellectual context. Tony Guida’s Paid to Be Central: Semantic Networks and the Cross-Section of Returns, presented at the CQF Institute conference on September 16, 2026, motivated the network-position baseline. Our narrow ontology projection is not a replication of his corpus or reported strategy. Centrality conventions follow Katz (1953); a practical mathematical reference is the NetworkX documentation. Figures and calculations: Lunartulip Lab, frozen AlphaMap Experiment 001 inputs.