A stock-analysis algorithm. Seven signal dimensions, 174 assets scanned every day. After six months of measurement, it delivered its verdict: don't listen to it. That's not a project failure. It's the most honest result it could produce. Camilo, a digital marketer in Valais and a complete beginner at investing, asks himself a simple question in January 2026: can an algorithm really beat the market? Rather than looking for the answer in a forum or an advert, he builds a tool to measure it himself, with real money: CHF 100 every week in a portfolio internal to OSOM Labs, held with Revolut, then every 14 days since August. Every order is decided and placed by hand. This is neither a service nor advice from OSOM Labs, and it isn't a bet: it's an experiment. Story updated on 24 September 2026. In the meantime, a full audit showed that the tool meant to tell the truth was getting things wrong too, and the algorithm became a coach. Both are covered below.
Seven dimensions, 174 assets, a check every evening
The algorithm analyses 174 assets across seven signal dimensions and several timeframes (multi-timeframe). A paper-trading system (simulated positions, no real money) lets it test its own decisions without risking anything. The whole thing runs unattended every evening at 10.30pm, with an automatic catch-up run in the morning if the previous evening's run failed.
Until the end of August, it sent a brief on Telegram every evening. Since 31 August 2026, it has worked in silence: only failures still trigger an alert. As of 24 September, 851 automated tests check that nothing breaks silently.
That's the algorithm. It took a few months to build. It isn't the point of this story.
The tribunal: three verdicts, in numbers
A signal that looks convincing on paper proves nothing. So rather than trusting its own results, the algorithm is built to judge itself: walk-forward validation (no information from the future leaking into the past), validation on assets and periods held back from tuning (out-of-sample), and a systematic comparison against the simplest strategy there is, buy and hold, real fees deducted every time.
Three independent measurements, run at different times, converge on the same conclusion.
First verdict: across the seven original signal dimensions, 0 out of 12 tested assets beat buy and hold. Pooled success rate: 39%.
Second verdict: the most promising candidate, a cross-sectional momentum strategy across several assets, is put through a stress test of 12 scenarios. It survives only 2. Rejected.
Third verdict, the most precise one: an empirical calibration of the composite score, run in July 2026 on five years of data across 12 assets, over a 20-day horizon. The table speaks for itself:
| Score band | Success rate | Sample |
|---|---|---|
| 0.0 – 0.2 | Success rate50% | Sample11,013 observations |
| 0.2 – 0.4 | Success rate47% | Sample2,522 observations |
| 0.4 – 0.6 | Success rate53% (inconclusive) | Sample17 observations |
| ≥ 0.6 | Success rate— | Sample0 observations in 5 years |
A score of 0.0 to 0.2 succeeds one time in two. So does a score of 0.4 to 0.6, but on 17 cases, no conclusion can be drawn. And nobody has ever seen a score above 0.6 in five years of market data. The signal says no more than a coin flip.
Direct consequence: the dashboard column that used to be called "Confidence" is renamed "Strength", with a plain-text warning attached. And the real money (CHF 100 per contribution, every week and then every 14 days since 9 August 2026) goes into buy and hold via scheduled instalments (DCA), not into the algorithm's signals. Active trading itself stays entirely in paper trading: zero real francs at stake.
In concrete terms, as of 24 September 2026, the plan has six lines and all six are held, after nine purchases. The target allocation is written in advance: 30% world ETF, 30% Nasdaq-100, then 10% each for Alphabet, semiconductors, European dividends and bitcoin. Each contribution goes to a single line, the one furthest from its target. No returns are published here: this isn't a performance being shown off. It's a discipline being documented.
The CRWD affair: a -71.8% that never existed
On 2 July, CRWD stock undergoes a 4-for-1 split in the paper portfolio. The system doesn't detect it. The next day, the dashboard shows a -71.8% loss on the position. A figure that sends a chill down your spine, until you check: the position is in fact profitable, up around +13%.
The choice, at that moment, isn't a minor one. You could let it slide, quietly patch it, or simply ignore a bug on a paper account; nobody would know. Instead, automatic split detection is added to the system, and the historical data is recalculated and retroactively corrected. The error stays documented in the history, not erased.
When your algorithm shows -71.8% and the truth is +13%, the question isn't "how do I fix the number"; it's "how many other numbers are we looking at without checking them".
The algorithm that refuses to declare victory
What follows is the most interesting part. Once the bug is fixed, the paper account used to explore the signals finds itself, as of 10 July, ahead of its own benchmark. Tempting to declare victory.
The system refuses to. Five weeks of data is statistically nothing: the significance thresholds are hard-coded, and they aren't met. No conclusion drawn from too small a sample, even when that small sample happens to be favourable.
What came next proved it right. As of 24 September, the same simulated account stands at -9.5%, while the Nasdaq-100 (QQQ) stands at -0.6% over the same period. The July lead proved nothing, just as the system had said. The simulated account dedicated to bitcoin did end up buying: it stands at +27.7%, against +29.8% for bitcoin simply bought and held over the same period. In both cases, buying and holding would have done better.
So the algorithm hasn't become useless. It has become what it should have been all along: an anti-FOMO discipline coach, a learning lab with no real money at stake, a safeguard that refuses to lie to itself, even when the truth would be flattering.
When the tool of truth gets it wrong
On 23 September 2026, the project goes through a full read-only audit: the algorithm, the platform, the strategy and the next purchases. Five audits, each followed by a counter-check that went back to the primary sources. The result: 67 findings. 49 confirmed, 16 qualified, 1 refuted and then reworded, 1 impossible to verify.
Two of its conclusions concern the tool meant to tell the truth.
The first is a currency error. The journal recorded the cost of purchases in dollars, and the dashboard converted it into francs at the day's rate, not at the rate on the day of purchase. The amount shown as invested was therefore around 2% higher than what had actually been paid, and the gain presented in francs was really a gain in dollars.
The second is simpler, and more awkward. Two purchases placed in the app had not been entered in the journal. For sixteen evenings, the coach therefore advised buying a line that had already been bought. Nothing allowed it to notice.
Both are fixed: the invested amount is now calculated at historical cost in francs, and the journal has been reconciled line by line with the account. As with the CRWD affair, the errors stay documented, not erased. A tool that tells the truth can get things wrong too; the only protection is a check that doesn't depend on it.
The process, never the timing
After the audit, the algorithm's role is redefined in writing. On real money, it applies rules written in advance: which line receives the next contribution, how much, and on which scheduled date. It never takes a view on when to enter the market, nor on a single stock that a signal might highlight. The decision and the order remain human.
The signals haven't disappeared, but they have lost their verbs. On the plan's lines, no more "buy" or "wait": just an "educational signal" with its score, and no colour nudging towards action.
And the coach has learnt to stay quiet. It would rather refuse to advise than advise wrongly. Two blocks stop it:
In both cases, the coach says why it is blocking and what needs to be done to unblock it. It is less comfortable than a recommendation every evening. It is the direct answer to the sixteen evenings of wrong advice.
The method: separating the one who produces from the one who judges
A word on how this system was built, because it follows the same logic the algorithm applies to the market.
Every change follows the same path: a written plan, an implementation, the test suite (851 tests as of 24 September), then a review by independent reviewers whose only job is to find the flaw. The first batch of fixes from the audit was reviewed by three of them. They still found seven defects, including one that could have led to buying the same line twice. Each defect got its own test.
The principle is the same everywhere: the one who produces is never the only one who judges. The algorithm doesn't trust its own signals until an independent backtest has validated them. The code doesn't go into service until an independent review has tested it. It's less spectacular than a demo that promises everything. It's also what makes it trustworthy.
Measuring before promising
OSOM Labs builds systems that measure before they promise. This internal project takes that to an extreme: an algorithm whose honesty is wired into the code rather than promised in a brochure, down to the "educational signal" label on the plan's lines, and the blocks that stop it advising on incomplete data.
In a market where everyone sells "the AI that predicts", this algorithm earned the right to say "I don't know", and then "I got it wrong", and all of it is documented. A target is never a guarantee. Here, the measurement runs every evening at 10.30pm, in silence; it isn't a slogan.
Reading frame
This story describes an educational project by OSOM Labs, carried out on a portfolio internal to the company. Decisions and orders are made and placed by hand. OSOM Labs offers no investment advice, no wealth management and no investment tool; the code is not distributed. The figures describe a system under test on a small portfolio: they are neither representative, nor reproducible, nor indicative.
Key takeaways
Three independent measurements converge: the algorithm's signals don't beat buy and hold (0/12, 2/12, ~50% success rate: a coin flip).
Once calibration proved the "Confidence" column measured nothing, it was renamed "Strength": the dashboard says what it knows, not what's reassuring.
A bug showed -71.8% on a position that was actually up +13% (an unhandled CRWD split): the data was corrected and the error documented, not erased.
The audit of 23 September 2026 raised 67 findings and showed that a dashboard presented a gain in dollars as a gain in francs: the tool of truth was getting things wrong too.
On real money, the coach applies a written plan (which line, how much, which date) and refuses to advise when its journal or its data are incomplete.
The most useful algorithm isn't the one that predicts. It's the one that stops you lying to yourself.
