Instagram
COPY CLIPBOARD
COPY FULL PAGE — PASTE INTO YOUR AI ↗
OPEN REQUEST ↗
MARKETS
LOADING MARKET DATA...
// ST-TRADER-001 — DATA OPERATIONS FOR MARKET OPERATORS
You have a thesis.
The dataset to
test it against
doesn't exist yet.
Bloomberg is the price layer. Synterminal is the physical market intelligence layer underneath it. Alternative datasets, normalized and provenance-anchored, built for your specific signal hypothesis — not for a generic subscriber base.
ST-INTELLIGENCE // SIGNAL MONITOR
● LIVE FEEDS --:--:--
// COMMODITY INTELLIGENCE — PHYSICAL LAYER SHA-256 ANCHORED
VERTICAL SIGNAL LEAD TIME STRENGTH
COPPER PROCUREMENT DEMAND ↑ ACCELERATING 6–18MO
HIGH
NAT GAS WEATHER BASIS CONTANGO +$0.18 3–7 DAYS
MED
RARE EARTH EXPORT CONTROLS ↓ SUPPLY SHOCK IMMEDIATE
CRIT
LITHIUM PROCUREMENT HEATMAP +357% DEMAND 12–24MO
HIGH
INVOICE Δ vs COMEX SPOT +$0.34/lb DRIFT LIVE
MED
// EQUITY & OPTIONS — ALTERNATIVE DATA LAYER RFC 3161 TIMESTAMPED
DATASET STATUS TYPE SIGNAL
GOV PROCUREMENT → SECTOR DEMAND NORMALIZED LEADING
LIVE
TARIFF IMPACT ON BOM COSTS BUILDING STRUCTURAL
MED
ENTITY SUPPLY CHAIN MAPPING ACTIVE ALT DATA
MED
// PROCUREMENT SIGNAL INTENSITY — 12MO ROLLING (ILLUSTRATIVE)
$2.8B
ALT DATA SPEND 2025
69%
NOT USING AI DATA OPT.
25–35%
DAY ON MANUAL ASSEMBLY
$27K
BLOOMBERG / SEAT / YR
// ST-TRADER-000 — THE UNDISCOVERED PUBLIC KNOWLEDGE PROBLEM
ABDUCTIVE REASONING · KNOWLEDGE GRAPHS · CROSS-DOMAIN SIGNAL
// THE SCIENCE BEHIND THE HUNCH
The connection exists
in the gap between
two datasets that have
never spoken to each other.
In 1986, Don Swanson discovered that fish oil treats Raynaud's syndrome — not in a laboratory, but by connecting two bodies of published literature that had never cited each other. Neither literature mentioned the other. The connection was real, public, and invisible. He called this undiscovered public knowledge — information that exists, is correct, and remains hidden because no one built the bridge between the two datasets that contained it. SOURCE: SWANSON LINKING / LITERATURE-BASED DISCOVERY — SPRINGER 2020
A
YOUR DATASET
KNOWN
B
THE BRIDGE NODE
UNDISCOVERED
C
THE SIGNAL
The formal name for the reasoning that produces a non-obvious hypothesis from incomplete data is abductive reasoning — inference to the best explanation, formalized by Charles Sanders Peirce. Deduction proves. Induction generalizes. Abduction guesses correctly from pattern and then tests the guess against data. Every non-obvious alpha thesis begins with abduction. The hypothesis arrives before the dataset that confirms it. The data confirms the direction — but the mechanism, the specific B node connecting A to C, is almost never what the trader predicted. That gap between the expected mechanism and the actual one is where the structural edge lives. SOURCE: DELLSÉN — ABDUCTIVE REASONING IN SCIENCE, CAMBRIDGE 2024
Synterminal's background is semantic web engineering and knowledge graph construction. The finding from that work: when heterogeneous datasets are connected through a shared ontology, the resulting graph exhibits emergent network properties — hub nodes, bridging concepts, multi-hop relationships that are invisible in any single dataset. The signal is not in the data. The signal is in the edge. Every engagement builds its own ontology from scratch, because no two hypotheses share the same structure of reality. The shared vocabulary between datasets is constructed for the specific thesis — not imported from a template. SOURCE: LLM-EMPOWERED KNOWLEDGE GRAPH CONSTRUCTION — ARXIV 2025 · ACM — CROSS-DOMAIN INSIGHT DISCOVERY VIA KNOWLEDGE NETWORKS 2026
If you have ever looked at two seemingly unrelated datasets and had a sense there was something there — that sense is abductive reasoning. It is how every non-obvious discovery begins across every scientific discipline. You do not need to know what the B node is. You need to know A and C. Synterminal finds B. The full force of forensic data engineering, semantic graph construction, and numerical pattern detection points at the hypothesis. What comes back is ground truth — provenance-anchored, queryable, reproducible — regardless of whether the signal holds or not. A confirmed null is also a finding.
// WHAT GROUND TRUTH MEANS ON EVERY ENGAGEMENT
PROVENANCE
Every record traceable to its source. SHA-256 hash at ingestion. RFC 3161 timestamp. The dataset cannot be disputed after the fact — the cryptographic record exists from the first record ingested.
QUERYABLE
Every dataset delivered as a queryable structure. Not a report. Not a PDF. A structured dataset your model can run against, your analyst can interrogate, and your compliance team can audit.
ONTOLOGY
Every engagement builds its own ontology. The shared vocabulary connecting disparate datasets is constructed from the specific entities, relationships, and time windows your hypothesis requires. No two are the same.
EMBEDDINGS
Semantic and numerical edges together. Vector embeddings capture conceptual similarity across datasets that share no common identifier. The B node that bridges A and C may not share a name with either — but it shares a semantic neighborhood.
// SIGNAL GRAPH — CROSS-DOMAIN INTELLIGENCE ST-GRAPH-001
Synterminal cross-domain signal graph
// ST-TRADER-007 — FREQUENTLY ASKED QUESTIONS
ALL TRADER TYPES · ALL ASSET CLASSES · ALL FIRM SIZES
I have a hunch about a connection between two datasets that seem unrelated. Is that enough to start?+
Yes. That hunch is the engagement. In knowledge discovery research, the hypothesis that two unrelated datasets share a hidden structural connection is called abductive reasoning — and it is formally the starting point of every non-obvious discovery. Swanson's ABC model shows that when A-B and B-C relationships are known, the A-C connection is a candidate signal. You identify A and C. Synterminal finds B. Describe what you're sensing. The engagement scopes from there.
Does this only work for commodity traders? I trade equities.+
The page goes deep on commodity and procurement intelligence because that is where Synterminal's proprietary signal work is most developed — but the data engineering capability applies to any asset class with any alternative dataset. A long/short equity thesis that requires sector procurement data, a volatility strategy that needs weather-correlated energy demand series, a credit thesis that requires supply chain concentration risk quantification — those are all the same engagement structure. The dataset changes. The methodology does not.
What if I don't know whether the data I need exists publicly?+
That is part of the scoping process — not a prerequisite for starting. Describe the hypothesis. Synterminal identifies what public datasets are relevant, which are machine-readable, and what engineering work is required to connect them. In many cases, the relevant data exists across three or four public sources that have never been normalized into a unified schema. The data exists. The clean dataset does not yet.
I already have Bloomberg and Refinitiv. What does Synterminal add?+
Bloomberg and Refinitiv are the price layer — real-time exchange data, news, standardized financial analytics. Synterminal is the physical and alternative intelligence layer underneath the price. Public procurement bid award flows. Supply chain concentration and chokepoint monitoring. Cross-domain signal construction connecting datasets that share no common schema or identifier. None of that is available on Bloomberg. Bloomberg's January 2026 introduction of alternative data entitlements confirms the market is moving in this direction — but custom dataset construction for a specific thesis is not something any terminal can deliver.
How is this different from buying a dataset on Neudata or Eagle Alpha?+
Neudata and Eagle Alpha list 7,000+ datasets from 4,200+ providers. Those are generic datasets built for a broad subscriber base — priced at $50,000–$200,000+ annually per dataset, regardless of whether the dataset fits your specific hypothesis. Synterminal builds a custom dataset for one specific thesis, normalized to your schema, delivered in a format your model can run against immediately. If the signal holds, the dataset becomes a live retainer feed. If it doesn't, the engagement cost a fraction of a vendor subscription that might not fit. The other difference: every Synterminal dataset is point-in-time correct, cryptographically provenance-anchored, and built with an ontology designed for your specific entities and relationships — not for the average subscriber.
What happens if the signal doesn't hold in the backtest?+
A confirmed null is a finding. You now know — with a provenance-anchored, reproducible dataset — that the A-C connection does not hold over the historical period tested. That is more valuable than intuition-based inaction. It also frees the research budget and attention to test the next hypothesis. In practice: the hypothesis almost always survives contact with the data. The direction is right. The mechanism — the B node — is often not what the trader predicted. That gap between expected mechanism and actual mechanism is where the most durable edges tend to live.
Is my hypothesis and data kept confidential?+
Every engagement is governed by a mutual NDA before any hypothesis or proprietary data is disclosed. The ontology built for your engagement, the custom dataset constructed, and any findings are yours. Synterminal does not aggregate or reuse client-specific datasets across engagements. The provenance architecture actually makes this enforceable — every record in your dataset carries a UUID and timestamp that is specific to your engagement, not a shared corpus.
I run a long/short equity strategy. I have a thesis on infrastructure buildout and copper-exposed equities. What can Synterminal build?+
The procurement heatmap is the primary dataset for this thesis. Public bid award data from federal infrastructure contracts, state DOT projects, utility commission procurement, and municipal public works — normalized by commodity category, geography, and quarter. When public agencies are accelerating electrical infrastructure procurement, that signal appears in bid award data 6–18 months before it appears in analyst coverage of copper miners or electrical equipment manufacturers. Synterminal builds the historical series going back to available data and delivers a live update protocol so the series stays current. Your model runs against a clean, point-in-time-correct dataset — not a CSV export that you have to reshape before it's usable.
I trade options and want alternative data around energy sector volatility. What's available?+
Several angles are constructable: NOAA weather station anomaly data normalized against natural gas front-month settlement, utility procurement bid calendars (which correlate with seasonal demand planning), EIA storage report history cross-referenced against weather deviation from seasonal norms, and supply chain disruption event chronology for energy infrastructure components. The combination — weather anomaly, procurement acceleration, and supply disruption mapped against options expiry windows — is a cross-domain dataset that no single vendor provides. The engagement defines which combination your volatility thesis requires.
Can Synterminal work with data I already have that needs normalization?+
Yes. If you have proprietary data — internal transaction records, vendor pricing history, customer behavior data, any structured or semi-structured source — Synterminal normalizes it: consistent schema, bi-temporal timestamps, UUID v5 entity identity, append-only WAL. The normalized proprietary data then connects to the public datasets through the shared ontology. That connection is where the cross-domain signal lives — your private A connected to a public C through a public B that you didn't know existed.
I'm doing fundamental equity research, not systematic. Is this relevant to me?+
AI eliminated the 3–4 hours of data assembly that precede fundamental analysis. What it did not do is normalize the alternative data that precedes the assembly. A fundamental analyst researching a supply chain–exposed industrial company needs the same procurement signal and supply chain concentration data that a systematic trader does — they just use it differently. The scoped research engagement format is designed for this: one specific question, one structured output, delivered in a format that goes directly into the research memo rather than requiring further engineering work.
I run a macro fund with positions across energy, metals, and agricultural commodities. What does a retainer engagement look like?+
A macro retainer covers multiple verticals simultaneously. Monthly deliverables are scoped per engagement — but a typical structure for a multi-commodity macro fund: procurement heatmap updated quarterly across your commodity exposure universe; supply chain chokepoint monitoring with flagging when a significant disruption event occurs in a watched vertical; one combination signal run per month across the cross-dataset layer; corpus integrity verification on all maintained datasets. The retainer scope is defined at engagement start and revised each quarter based on where the research is focusing. No deliverable is a report. Everything is a queryable, structured dataset your team can run against directly.
I trade physical copper. How does procurement signal data improve my positioning?+
Public procurement bid awards for electrical infrastructure — wire, cable, conduit, switchgear — are filed 6–18 months before the physical demand materializes on the exchange. When the state of California awards $340M in electrical infrastructure contracts in Q1, that copper demand does not appear in COMEX open interest until Q3 at the earliest. The bid award data is public, machine-readable, and almost entirely unstructured as a time series. Synterminal builds the structured series. The physical trader who has that series has a demand signal that the exchange price has not yet incorporated.
What's the Swanson Linking concept and why does it matter for my commodity research?+
Swanson Linking is the formal name for cross-domain discovery — finding an A-C connection between two datasets that have never been connected, using a B dataset as the bridge. In commodity markets: A is your commodity price history. C is a supply or demand event you're trying to predict. B is a dataset — procurement flows, weather anomaly, entity supply chain map — that connects A to C through a relationship that wasn't previously visible because the B dataset was never structured as a time series. Synterminal's ontology layer is the mechanism that makes B visible. The ontology defines the shared vocabulary that allows A and B and C to be queried together as a single dataset.
I trade agricultural commodities. Can you build weather-correlated demand signal data?+
NOAA weather station data, USDA crop condition reports, and public procurement bid awards for food and agricultural commodity inputs can all be normalized into a unified time series. The combination — weather anomaly in a key production region, USDA condition rating deviation from seasonal norm, and procurement acceleration or deceleration in downstream food processing — is a cross-domain signal that no single data vendor provides in structured form. The engagement defines which crop, which geography, and which time window the thesis requires. The dataset is built from those parameters.
How does a scoped engagement actually start?+
A request through the contact form describes the hypothesis — the A and C you're trying to connect, the asset class, and the time window that matters. Synterminal responds with a scope document: what public data sources are available, what engineering work is required to normalize them, what the output format will be, and what the timeline and cost look like. The scope document is free. The engagement begins when the scope is agreed. No vendor lock-in, no annual subscription commitment at the scoping stage.
How long does a scoped engagement take?+
A single-dataset normalization engagement — one public source, one schema, historical series back to available data — typically delivers in 1–2 weeks. A cross-domain combination engagement involving 3–5 sources and a custom ontology build is typically 3–5 weeks depending on source complexity. A backtestable historical dataset construction engagement for a multi-factor signal depends on how far back the source data goes and how many entity resolution passes are required. Every scope document includes a timeline estimate. The timeline is fixed at scope agreement, not billed hourly.
I'm a solo trader running my own capital. Is this only for funds?+
Solo operators are a primary client segment. A full-time data engineer is not justifiable at any AUM level below $50–100M. A scoped engagement that delivers exactly the dataset the thesis requires — once, for a flat fee — is exactly what a solo systematic trader needs. The retainer model scales to whatever the ongoing data maintenance requirement actually is: one dataset updated monthly, or five datasets with a combination signal run each quarter. The scope defines the cost. There is no minimum AUM.
What markets do you cover? US only?+
Public procurement data is available across the US federal government, all 50 states, and major municipal jurisdictions. NOAA and USDA data covers the US. International procurement data is available through EU public procurement portals, UK government contracts finder, and several other national procurement systems — though coverage depth varies by jurisdiction. HTS trade flow data covers all US import/export customs declarations and can be cross-referenced against international counterpart data. The engagement scopes to the geography the thesis requires and confirms data availability before the engagement begins.
What format is the output delivered in?+
Every dataset is delivered as a structured file in the format the client's infrastructure requires — CSV, Parquet, JSON-L, or SQL-compatible schema. The dataset ships with full documentation: schema definition, field descriptions, source provenance for every record, the ontology specification used to connect heterogeneous sources, and the hash verification file for corpus integrity checking. It is not a report. It is a dataset your model can run against, your analyst can query, and your compliance team can audit to any record.
What does the retainer relationship actually look like month to month?+
The retainer delivers five things each month: one new dataset normalized or one existing feed updated when a source changes schema; one combination signal run per the current thesis; corpus integrity verification across all maintained datasets; one anomaly or opportunity flagged when the data shows something the thesis predicts or contradicts; and a full audit trail maintained on every dataset. The scope is fixed at retainer start and reviewed quarterly. You are not billed hourly and you do not manage a team. You receive structured intelligence on a schedule and the data infrastructure stays clean between engagements.
What is the difference between a scoped engagement and a retainer?+
A scoped engagement is a defined deliverable with a fixed cost and a fixed timeline — one dataset built, one hypothesis tested. A retainer is an ongoing data operations relationship — the dataset stays current, new sources are added as the thesis evolves, and the cross-domain analysis runs on a schedule. The typical path is: scoped engagement first, retainer if the signal holds and the thesis needs ongoing data maintenance. The scoped engagement cost is not a deposit toward the retainer — it is the evaluation cost for the hypothesis.
How is the retainer priced?+
Retainer pricing is scoped per engagement — not a published rate card, because no two engagements cover the same datasets, the same combination complexity, or the same update frequency. The market benchmark for ongoing data science advisory retainers in 2026 runs $3,000–$15,000 per month depending on scope and hours. Synterminal's retainer pricing reflects the actual scope: light maintenance on a single dataset is at the low end; active cross-vertical combination analysis with multiple live feeds is at the high end. The scope document specifies the monthly cost before the retainer begins.
Can I cancel the retainer?+
Retainer engagements operate on a rolling monthly basis with 30-day notice to cancel. There is no annual commitment required. The datasets constructed during the retainer are yours — they leave with you if the retainer ends. The ontology specification and hash verification files are included in the final deliverable package. The only thing that stops is the monthly update and combination analysis work.
I'm an RIA with commodity exposure in client portfolios. Is a retainer the right structure for ongoing intelligence?+
For an RIA or wealth manager with ongoing commodity or real asset exposure in client portfolios, a retainer typically covers: a monthly intelligence briefing on the physical market underlying the position — supply chain conditions, procurement demand signal update, any significant disruption or concentration event in the relevant commodity — delivered as a structured dataset and summary that goes directly into the investment committee presentation. The retainer replaces the time a portfolio manager or analyst would spend manually aggregating that intelligence from disparate sources, which typically consumes 3–4 hours per week at minimum.
// MARKET CONTEXT — THE GAP SYNTERMINAL OPERATES IN
ALL ASSERTIONS SOURCED
17%
Year-over-year growth in alternative data spend in 2025 — market hit $2.8B, on track for $168B by 2034
69%
Buy-side firms that have NOT adopted AI-processed data to optimize investment strategies — the gap is the next two years of work
25–35%
Of active trading day spent on manual data assembly at mid-market metals desks — pulling feeds, reconciling settlements, translating price moves to P&L views
52%
Of boutique asset managers planning to outsource key research and data functions in the next 12–24 months — fee compression making in-house research teams unjustifiable
// ST-TRADER-002 — BLOOMBERG VS SYNTERMINAL
THE PRICE LAYER · THE PHYSICAL INTELLIGENCE LAYER
// THE DISTINCTION — MADE ONCE
Bloomberg delivers the price. Synterminal builds the dataset that explains why the price moved before it moves.
Bloomberg Terminal at $24,000–$27,000 per seat per year delivers real-time price data, news, analytics on public markets, company financials, and standardized macro feeds. It is the best instrument ever built for reading the market as it exists right now. What it does not do: ingest a custom alternative dataset, normalize a non-standard source, run pattern recognition against physical market flows, or build a procurement demand signal from 6 million public bid awards. Bloomberg is the price layer. The physical intelligence layer underneath it — the one that explains why price is moving — requires a different kind of infrastructure. SOURCE: CATALAYER — BLOOMBERG TERMINAL GUIDE 2026
The advantage in alternative data no longer comes from accessing it — 7,000+ datasets exist on Neudata alone. The advantage comes from extracting and operationalizing the right signal at scale. The engineering task of structuring data pre-algorithm is the documented bottleneck. Data sources are fragmented. Schemas are inconsistent. Timestamps are not point-in-time correct. The model cannot run on dirty data. That engineering problem is the engagement. SOURCE: COGNITE — FIVE DATA CHALLENGES IN COMMODITY TRADING
// CAPABILITY COMPARISON
CAPABILITY BLOOMBERG SYNTERMINAL
Real-time exchange price Best in class Not the product
Public procurement signal Not available Core vertical
Custom alt dataset build Not available Primary service
Invoice-level pricing delta Not available Ledger pipeline
Schema normalization Your problem Every engagement
Point-in-time correctness Not guaranteed Bi-temporal, day one
Combination asymmetry No cross-vertical Core thesis
Cryptographic provenance Not applicable SHA-256 / RFC 3161
Per-seat cost $24–27K/yr/seat Scoped engagement
// ST-TRADER-003 — SIGNAL VERTICALS
COMBINATION ASYMMETRY — NO SINGLE VERTICAL IS THE EDGE
// THE COMBINATION THESIS
No single dataset is the edge. The signal lives in the combination no one has assembled.
Physical commodities are the biggest diversification play in 2026. Larger firms and start-ups are hunting for alpha that quant approaches cannot easily access. The datasets that describe the physical market — procurement flows, supply chain concentration, mine production, invoice-level pricing, weather correlation to energy demand — are all public. None of them are clean. None of them talk to each other. The combination, structured correctly, is where the edge lives. SOURCE: WITH INTELLIGENCE / YAHOO FINANCE — PHYSICAL COMMODITIES 2026
Each new data vertical multiplies the edge of every existing one. Procurement heatmap alone tells you where demand is building. Combined with NOAA weather anomalies — it tells you whether that demand is being met. Combined with invoice-level pricing — it tells you whether the market has priced the supply constraint into procurement contracts yet. That combination is not available anywhere at any price. Synterminal builds it record by record.
// ACTIVE SIGNAL VERTICALS
PROCUREMENT HEATMAP 6–18MO LEAD
PHYSICAL SUPPLY CHAIN CHOKEPOINT MONITOR
WEATHER × ENERGY NOAA ANCHORED
INVOICE Δ vs SPOT REAL PRICING LAYER
HTS TRADE FLOWS PORT + CUSTOMS
ENTITY SUPPLY CHAIN WHO OWNS THE MINE
// ARCHITECTURE — EVERY RECORD
HASH SHA-256
TIMESTAMP RFC 3161
IDENTITY UUID v5
WRITE LOG APPEND-ONLY WAL
TIME MODEL BI-TEMPORAL
// ST-TRADER-004 — SERVICE VERTICALS
COMMODITY · EQUITY · OPTIONS · MACRO — ALL TRADER TYPES
// ST-TR-V01 — ALTERNATIVE DATASET CONSTRUCTION
The hypothesis exists. The clean dataset to run it against doesn't.
The engineering task of structuring data pre-algorithm is the documented bottleneck for systematic commodity strategies. Synterminal does the build so the model can run.
Any source that produces structured or semi-structured data is ingestion material: public procurement databases, NOAA weather station records, HTS trade flow files, port freight data, EIA energy reports, USDA agricultural surveys, Secretary of State entity records, court filings. Synterminal normalizes the schema, enforces consistent timestamp formats, applies UUID v5 identity across entities, hashes every record at ingestion, and delivers a clean dataset with full provenance documentation. The model receives a dataset where every row is traceable, every timestamp is point-in-time correct, and the integrity of the corpus is verifiable against a Merkle tree. That is the difference between a model that backtests on clean data and one that backtests on noise.
// ST-TR-V02 — PROCUREMENT HEATMAP AS DEMAND SIGNAL
Public procurement bid awards are machine-readable, publicly available, and almost entirely unstructured as an investment dataset.
When Southern California public agencies accelerate copper wire procurement ahead of a grid modernization mandate, that signal appears in public bid data 6–18 months before it appears in COMEX open interest. Bloomberg does not have it structured. No vendor does.
The commodity ontology layer maps public bid line items to commodity categories, geographic regions, and time windows. Copper wire bid awards in California jurisdictions. Steel procurement across federal infrastructure contracts. Lithium and rare earth components in Department of Defense supply agreements. The output is a structured time series: procurement volume by commodity, by region, by quarter. Each record carries the bid number, the awarding agency, the awarded amount, the commodity classification, and the timestamp of award. That time series is a leading demand indicator — it reflects real purchasing decisions made by real buyers before the commodity market prices the demand in. For an equity trader with a thesis on copper miners or electrical grid suppliers: this is the dataset that tests it.
// ST-TR-V03 — PHYSICAL SUPPLY CHAIN INTELLIGENCE
Mine production disruptions. Refining concentration shifts. Export control chronology mapped to price action. All public. None of it structured.
Smaller physical traders trail leading trading houses by more than 20%. The gap is technology and data infrastructure, not trading judgment.
Chokepoint monitoring: Suez and Panama Canal freight conditions, refining capacity utilization by country and commodity, mine production event chronology (strikes, slides, weather shutdowns, regulatory suspensions) cross-referenced against LME and COMEX price action in the same window. For copper: IEA projects a 25% supply deficit through 2035. The specific production events that feed into that deficit happen at named mines on datable days — and they are public record. Synterminal structures that chronology, timestamps it correctly, and delivers it as a dataset against which a price model can be run. For a commodity trader running physical copper, natural gas, or agricultural positions: this is the supply side intelligence that sits underneath the futures curve.
// ST-TR-V04 — EQUITY AND OPTIONS RESEARCH SUPPORT
AI eliminated the 3–4 hours of data assembly before analysis. It did not normalize the data. That is still manual at most desks.
52% of boutique asset managers plan to outsource key research functions. The first thing to outsource is the data that precedes the research.
For equity and options traders: sector-specific alternative datasets built to a hypothesis. A long/short equity fund with a thesis on electrical grid infrastructure buildout — Synterminal builds the procurement demand dataset that tracks federal, state, and municipal bid awards for electrical components across the target geography. An options trader running a volatility strategy around energy earnings — Synterminal structures the NOAA temperature anomaly dataset, the nat gas storage report history, and the utility procurement bid calendar into a unified time series. For any thesis that requires alternative data: a scoped engagement produces a backtestable historical dataset and a live feed protocol. The quant team runs the backtest. If the signal holds, the dataset becomes a retainer deliverable. If it doesn't, the engagement cost a fraction of a Neudata vendor subscription that might not fit the hypothesis.
// ST-TR-V05 — SCHEMA NORMALIZATION AND FEED COHERENCE
Multiple data subscriptions. Different schemas. Different timestamp formats. Different unit conventions. The feeds don't talk to each other.
64% of ETRM/CTRM systems still don't support all their processes. Reporting is the first visible bottleneck — teams export manually and rebuild in spreadsheets what the platform should handle automatically.
Exchange feeds, weather APIs, government data portals, internal trade records, and alternative data subscriptions each arrive in their own format with their own timestamp convention and their own entity naming schema. Making them interoperable — so a query across all of them returns a coherent result — is a data engineering problem that consumes analyst hours at every desk that hasn't solved it. Synterminal normalizes the stack: common schema across sources, bi-temporal timestamps enforced uniformly, UUID v5 identity so the same entity resolves to the same record across every source, append-only WAL so every change is logged and recoverable. The feeds stop being separate silos and start behaving as a unified dataset. Every downstream query runs against the same clean foundation.
// ST-TR-V06 — BACKTESTABLE HISTORICAL DATASET CONSTRUCTION
Before a fund evaluates whether a signal works, they need a clean historical dataset going back far enough to be statistically meaningful.
Backtesting alternative data is the documented bottleneck in dataset evaluation — opaque, unique to each fund, and inefficient. The dataset has to exist first. Synterminal builds it.
The quant process requires a historical baseline before any live signal evaluation begins. For any alternative dataset — procurement heatmap, supply chain disruption chronology, invoice-level pricing delta, weather-energy correlation — Synterminal constructs the historical version: every record normalized, provenance-documented, point-in-time correct back to the available source history. The fund receives a dataset against which their own backtest framework can run without modification to their existing infrastructure. If the signal holds over the historical period, the dataset converts to a live feed under a retainer arrangement. If it doesn't, the engagement cost a fraction of the minimum vendor contract that would have delivered the same data in a format incompatible with the fund's specific hypothesis. The evaluation cost is the scoped engagement. The production cost is the retainer.
// ST-TRADER-005 — WHO THIS IS FOR
FIVE SEGMENTS · SAME INFRASTRUCTURE · DIFFERENT THESIS
// SOLO SYSTEMATIC
Signal thesis. Model framework. No data engineering hire.
Running own capital or a small fund. The model is built. The dataset that feeds it is fragmented or doesn't exist in clean form. A full-time data engineer is not justifiable at this AUM. A scoped engagement that delivers exactly the dataset the thesis requires is.
Custom dataset build
Historical backtestable series
Live feed on retainer
// PHYSICAL COMMODITY DESK
Trading copper, nat gas, agricultural. Physical knowledge. Data assembly eating the day.
Mid-size trading desk with physical market expertise and Bloomberg for price. The supply chain intelligence layer underneath the price is manual — pulled from multiple sources, reconciled by hand, 25–35% of active trading time consumed before a single trade decision is made.
Procurement demand signal
Supply chain chokepoint monitor
Schema normalization
// EQUITY / OPTIONS FUND
Long/short or volatility strategy. Sector thesis needs a dataset that doesn't exist on any vendor platform.
Boutique equity or options fund with a specific sector thesis — infrastructure, energy transition, commodity-exposed industrials. The thesis requires alternative data that isn't available as a standard subscription. The data exists publicly. It isn't structured.
Sector procurement dataset
Thesis-specific alt data build
Research-ready delivery
// RIA / WEALTH MANAGER
Commodity or macro exposure. Real client money. No data scientist on staff.
Sub-$1B AUM, 1–10 people, managing concentrated positions in commodity-linked equities, real assets, or macro instruments. They have Bloomberg or an alternative stack. They do not have someone to answer a specific physical market intelligence question.
Monthly intelligence briefing
Procurement demand update
Hypothesis-specific research
// FAMILY OFFICE
Real asset exposure. Fastest-moving LP cohort in alternatives. Intelligence to match the position.
Gold, commodities, and infrastructure are top-five alternative allocations for family offices in 2025. 67% increasing alternatives exposure in 2026. They need physical market intelligence on the assets they hold. They will pay for a retainer if the intelligence is specific, sourced, and actionable.
Physical commodity intelligence
Supply chain concentration risk
Ongoing retainer relationship
// ST-TRADER-006 — THE RETAINER MODEL
WHAT THE ONGOING RELATIONSHIP LOOKS LIKE
// THE DATA OPERATOR ON RETAINER
Every serious desk has this person. At large funds they hire them full-time. At small funds they go without — until now.
The function is the same regardless of what the fund calls the role: quant researcher, data engineer, research analyst. Someone who keeps the data infrastructure clean, builds the next dataset when the thesis evolves, normalizes the new source when it arrives, flags anomalies when a schema changes, and delivers research-ready output on a schedule. Synterminal is that function, scoped to exactly what the engagement requires. SOURCE: INTERVIEW QUERY — DATA SCIENCE CONSULTANT RATES 2026
// MONTHLY RETAINER — WHAT IS DELIVERED
STANDARD SCOPE
01
One new dataset normalized or one existing feed updated. New source arrives — Synterminal ingests it, normalizes schema, enforces bi-temporal timestamps, delivers clean. Existing feed changes schema — Synterminal catches it before the model breaks.
02
One combination signal run per month. The thesis evolves — new question, new cross-vertical analysis. Procurement × weather × price. Invoice delta × supply disruption chronology. Whatever the hypothesis, one structured output per month.
03
Corpus integrity verification. SHA-256 hash check on every existing dataset. RFC 3161 timestamps verified. Any upstream source change flagged before it propagates into the model as silent data corruption.
04
One anomaly or opportunity flagged. Procurement acceleration in a target commodity. Supply disruption event cross-referenced against futures curve structure. Entity change in a watched supply chain node. When the data shows something, you hear about it.
05
Full audit trail maintained. Append-only WAL. Every change logged. Every version of every dataset recoverable. The record of what the data said at any point in the past is always available — which is what a backtesting framework requires and what a compliance review demands.
// MARKET CONTEXT — WHAT THIS WORK COSTS ELSEWHERE
Full-time junior quant researcher
$150–250K/yr
Bloomberg Terminal per seat
$24–27K/yr
Data science retainer — standard
Neudata vendor dataset — entry
$50–200K/yr
Synterminal — scoped engagement
CONTACT FOR SCOPE
FLAT-RATE OR RETAINER
Describe the thesis. Synterminal scopes the dataset and the engagement structure — one-time build or ongoing retainer.
research@synterminal.com
You have a thesis.
We build the
dataset.
Physical commodity traders, systematic funds, equity and options desks, RIAs, family offices. If the edge lives in data that exists publicly but isn't structured — that is the engagement. One scoped build. One decision on whether the signal holds. One retainer if it does.
// OPEN A REQUEST

Describe the thesis, the asset class, and what dataset you need that you can't currently get in clean form. Synterminal scopes the engagement. No vendor lock-in. No annual subscription to a dataset that doesn't fit the hypothesis.

research@synterminal.com
Bring the
ugly problem.
If a pricing question, data mess, supplier problem, monitoring task, or operational mystery has been sitting untouched — that is the job. Synterminal investigates what others don't have the infrastructure to find.
// OPEN A REQUEST

Direct intake for difficult information problems. Physical markets, procurement, pricing, supply chain, entity resolution, litigation support. Describe the problem. We scope the engagement.

research@synterminal.com