The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →S&P Global says it expanded RiskGauge coverage of U.S. private small and midsize enterprises (SMEs) from about 2 million to roughly 10 million—five times as many companies—by combining multi-layer website crawling, text extraction, ensemble machine-learning models, anonymized third-party data and Snowflake-based processing. The increase is a coverage claim, not evidence of a fivefold improvement in predictive accuracy.
The system turns public web signals into firmographic inputs for risk scores. It can help lenders, insurers, procurement teams and investors assess companies with limited conventional disclosure, but it does not replace financial statements, payment records or human review.
Why private SMEs are difficult to assess
Public companies usually provide standardized financial disclosures. Many private SMEs do not publish quarterly accounts or investor-grade reporting, so a lender or large buyer may lack basic evidence about a company’s existence, activity, sector, location, scale and momentum.
The problem is often coverage rather than total invisibility. A legitimate business may appear in registries, local news, social profiles, filings or a website, but those sources are fragmented and inconsistent. S&P told VentureBeat that its target was approximately 10 million active U.S. private SMEs, excluding sole proprietorships, compared with about 2 million previously covered. Those figures are reported estimates, not an independently audited count of every U.S. SME.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
What the fivefold expansion means
RiskGauge is designed to combine financial, business and market-risk information into scores, reports, monitoring and comparisons. S&P describes a 1-to-100 scale in which 1 is the highest credit risk and 100 the lowest. Its current RiskGauge Desktop materials position the product for customer, supplier and counterparty analysis.
The reported fivefold result primarily means more companies receive a usable record and score. It does not establish five times more data per company, five times better default prediction or equal data depth for all 10 million businesses.
“Deep web” in this implementation
Here, “deep” mainly describes crawling several layers inside a company’s domain rather than stopping at the home page. The interview refers to landing pages, contact pages, company descriptions, news and announcement pages and other text-bearing pages, along with anonymized third-party datasets.
That is different from claiming access to password-protected, paywalled or dark-web sources. The public description does not establish that restricted databases were accessed.
Free tools Windows power users keep installed
One-click scans. No signup required.
The pipeline from websites to risk scores
Company domains
↓
Crawlers and scrapers
↓
Preprocessing and text extraction
↓
Data-mining algorithms
↓
Curation and validation
↓
Firmographic drivers
↓
RiskGauge scoring
↓
Reports, monitoring, APIs and Snowflake delivery
1. Crawling multiple URL layers
Crawlers discover and fetch pages within company domains. The approach avoids relying entirely on rigid sitemaps or site-specific rules because websites vary widely. S&P discussed processing information from more than 200 million websites or pages and several terabytes of website data; the interview uses both terms, so the precise unit is not clear.
Rank #2
- Wiley
- Language: english
- Book - storytelling with data: a data visualization guide for business professionals
2. Preprocessing the content
The reported preprocessing layer removes material the models do not need, particularly HTML tags, JavaScript, TypeScript and other code, while retaining human-readable text. Normalizing pages this way makes it easier to process many unrelated site designs.
This does not prove that every structured element, image, PDF or table is discarded. The public account does not provide a complete retained-data schema.
3. Mining and curating signals
Specialized algorithms extract attributes such as company name, business description, sector, location and evidence of operating activity. Curators and validation logic then prepare those outputs as inputs to RiskGauge.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute4. Scoring the company
The web-derived variables become one evidence layer alongside financial risk, business risk, market risk, business-credit information, historical performance, key developments and peer comparisons. A website is not itself a credit score.
Where ensemble learning fits
S&P reported using multiple algorithms that examine different parts of a company’s pages and “vote” on conclusions. The models reportedly help identify company attributes and detect sentiment or polarity in announcements.
Rank #3
Ensembling can make a system less dependent on one classifier’s weaknesses. Agreement among models may also act as a quality signal when wording and page layouts differ. However, S&P has not publicly disclosed the model families, weights, thresholds, calibration method or evaluation statistics. The description does not establish whether the implementation uses random forests, boosting, stacking, neural models or a hybrid.
Snowflake’s role
According to the interview, Snowflake’s warehouse and Snowpark Container Services sit in the preprocessing, mining and curation stages. That provides a scalable environment for large text volumes, repeated transformations, custom algorithms and downstream integration.
Recommended Free Tools
Snowflake is infrastructure, not the proprietary RiskGauge methodology. The fivefold coverage result reflects the combined system: discovery, extraction, model design, third-party data, operations and product integration. Snowflake offers consumption-based pricing across Standard, Enterprise and Business Critical editions; costs depend on cloud, region, edition and workload. See the pricing page and service-consumption table.
How records stay current
S&P described weekly scans, with downstream updates triggered by detected changes rather than rewriting every record every week:
- A hash is created for a previously captured landing page.
- A later crawl creates a new hash.
- A matching hash indicates no detected content change.
- A mismatch triggers further processing or an update.
Hashing controls compute and reduces unnecessary model runs, but it detects content change rather than truth. A redesign can create a change event without a meaningful business development; a dynamic page can create repeated false positives; and an unchanged page can coexist with deteriorating finances. Website-update frequency is therefore a heuristic for activity, not proof of solvency.
Rank #4
The accuracy-versus-speed trade-off
S&P said some algorithms had strong accuracy, precision and recall but were too expensive to run at the required scale. Production choices therefore balance:
- Precision against recall and false-positive risk.
- Model sophistication against latency and compute cost.
- Crawl depth against coverage and expense.
- Freshness against the number of rescoring jobs.
- Generalization against the diversity of websites and sectors.
The practical lesson is that the best laboratory model may not be the best model for hundreds of millions of pages. Avoiding hard-coded rules and robotic process automation can improve resilience, but it also makes entity resolution and quality control harder.
What web signals can—and cannot—tell a lender
| Signal type | Useful indication | What it does not prove |
|---|---|---|
| Company name, location and description | Identity and basic firmographics | Legal ownership or financial strength |
| Service and sector language | Business activity and industry classification | Revenue, margins or market share |
| Announcements and sentiment | Potential developments or momentum | Liquidity or repayment capacity |
| Recent website updates | A possible activity indicator | Continued operation or solvency |
| Rich online presence | More extractable evidence | Lower default probability |
Web data can expand the evidence available for thin-file companies. It cannot substitute for cash flow, debt, liens, payment history, audited accounts or verified ownership. Marketing claims may be inflated, copied or maintained by an agency after a business has stopped trading.
Failure modes buyers should test
- Stale or thin websites: legitimate firms may have little online presence, while closed firms may leave sites online.
- Entity confusion: subsidiaries, franchises, trade names and similarly named companies can be conflated.
- Template contamination: website-builder boilerplate can make unrelated companies look alike.
- Dynamic or blocked pages: client-rendered content, rate limits and bot defenses can reduce coverage.
- Sector and language bias: digital-first or English-language businesses may generate richer signals.
- Proxy bias: web activity can reflect marketing budget, size or digitization rather than creditworthiness.
- Concept drift: language, designs and business practices change over time.
- Adversarial content: companies could optimize pages to influence automated classifiers.
Ensemble agreement does not remove shared bias: several models can reach the same wrong conclusion from the same misleading page.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Governance and responsible use
The public account does not answer several questions that a regulated or high-stakes deployment must address:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
- Whether crawlers follow robots.txt and site terms in every jurisdiction.
- How personal information and employee data are filtered.
- How source URLs, timestamps, extraction methods, confidence and model versions are retained.
- How duplicate entities and parent-subsidiary relationships are resolved.
- How a company can correct an error or challenge an assessment.
- Whether scores support analysts or trigger automated decisions.
- How legal events, sanctions, adverse news and politically exposed persons are handled.
S&P reported no human in the initial ensemble decision process. That does not establish a human-free customer workflow; exception handling, complaints, model governance and material credit decisions may still require people.
Buy RiskGauge, use an API or build the stack?
| Option | Best fit | Main trade-off |
|---|---|---|
| RiskGauge Desktop | Enterprises managing large customer, supplier or counterparty portfolios | Packaged scores and monitoring require a demo and commercial agreement; no public self-serve price is shown on the reviewed page. |
| Universal Coverage API | Teams embedding entity matching, coverage or RiskGauge reports in underwriting or procurement software | Integration control, but account-based commercial engagement; no public price was displayed on the reviewed Marketplace page. |
| Snowflake-based build | Organizations with data-engineering capacity and proprietary data to combine | Elastic processing and governance, but consumption costs and operating complexity. |
| Complete custom crawler and model stack | Very large-volume users with a defensible need for proprietary features | Maximum control requires continual work on compliance, entity resolution, lineage, monitoring, quality and dispute processes. |
A custom build also needs crawl politeness, rate limiting, browser rendering where necessary, source retention, model monitoring and cost controls. Buying finished intelligence is usually simpler when the goal is immediate portfolio decisions; building makes more sense when proprietary data, workflow integration or unusual coverage requirements justify the ongoing expense.
What remains undisclosed
The available public material does not provide AUC, KS, Brier scores, calibration statistics, default-prediction lift, false-positive rates, out-of-time validation, sector-by-sector results or a comparison of web-only with web-plus-financial models. It therefore supports a description of architecture and coverage expansion, not a claim of predictive superiority.
For buyers, the most important diligence requests are model-performance evidence, coverage definitions, provenance and timestamps, correction procedures, score explanations, refresh behavior and the role of human analysts in consequential decisions.
The Bottom Line
S&P’s project shows how industrial-scale crawling and machine learning can bring millions of thin-file SMEs into a risk workflow. Its durable lesson is architectural: public web signals can broaden coverage when they are extracted, validated and monitored at scale. They remain proxies, however, and should complement—not replace—verified financial and payment evidence.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




