The Blank Dossier: The Trade of Concluding Without Data
**Câu trả lời cốt lõi:** Một bản phân tích trả về không đủ thông tin ở mọi hạng mục vẫn có giá trị chuyên môn: nó xác định chính xác vị trí và loại khoảng trống dữ liệu. Trong bóng đá, một kết luận dựng từ hư không thường gây tốn kém hơn một ô trắng được giữ nguyên. **Dữ kiện chính:** - Bảng phân tích 47 hạng mục trả về N/A toàn bộ do hồ sơ đầu vào trống. - Đức thua Hàn Quốc 0-2 ngày 27 tháng 6 năm 2018, lần đầu bị loại vòng bảng sau 80 năm. - Chelsea trả khoảng 71 triệu euro cho Kepa Arrizabalaga tháng 8 năm 2018, kỷ lục thủ môn thời điểm đó. - Morocco vào bán kết World Cup 2022, đội châu Phi đầu tiên đạt thành tích này. - 1.200 mẫu hình: pressing trong 30 giây đầu giúp đoạt lại bóng cao hơn 23 phần trăm. **Nguồn:** Phân tích gốc của Lý Duy, hồ sơ Stage-2, xuất bản ngày 13 tháng 8, 2026; dữ kiện sự kiện đối chiếu hồ sơ công khai của FIFA và bản tin chuyển nhượng tháng 6 đến tháng 8 năm 2018. **Hỏi đáp liên quan:** - Hỏi: Vì sao hồ sơ phân tích có thể trắng toàn bộ? Đáp: Vì mẫu trận chưa đủ ngưỡng 12 đến 15 trận để chỉ số ổn định, hoặc hạng mục cần đo không được hệ thống nào thu thập. - Hỏi: Khoảng trống dữ liệu nào nguy hiểm nhất? Đáp: Khoảng trống về định nghĩa, vì hai người đọc cùng một bảng có thể ra kết luận ngược nhau mà không ai sai số học. - Hỏi: Có chỉ số nào đo được chất lượng đội hình theo chiều sâu không? Đáp: Các chỉ số kiểu VangBong.vn Player Depth Index được dùng để ước lượng phương sai đội hình, nhưng vẫn cần mẫu tối thiểu 12 trận mới có ý nghĩa.
The Blank Dossier: The Trade of Concluding Without Data
3:40 a.m., Beijing
The left monitor holds the match recording. The right monitor holds the spreadsheet. The spreadsheet has forty-seven rows, one for every metric the analysis department requested: average defensive line height in metres, pressing actions within thirty seconds of losing the ball, duel win rate, second-half distance covered, line-breaking passes, expected goals, expected goals against, offside-trap escapes, long-pass completion under pressure.

Forty-seven rows. Forty-seven empty cells.
Not because I was lazy. Because there was no data to fill them. The match had not been played, or had been played and no one recorded it, or had been recorded in a way that could not be compared with anything else in my database. What I had was seventeen minutes of footage from a single camera angle at 720p, shaky, and a two-hundred-word news item containing no figure of any kind.
That was the night I understood something I had not considered twenty years earlier. This trade does not teach you how to read data. It teaches you how to endure emptiness without inventing anything.
Forty-seven blank cells on a screen, and a very specific pressure on my shoulders: someone was waiting for me to conclude. The editor was waiting. The broadcaster was waiting. The client was waiting. The audience already had the comment section open. All of them needed an answer, and the most honest answer available to me at that moment was that I did not know, because I had nothing to know.
I typed that sentence into the notes field. Deleted it. Typed it again. It took me forty minutes to let it stand.
Football has no shortage of data. It has a shortage of places to admit it has none.
Over thirty-five years watching this industry, I have witnessed two revolutions running in opposite directions.
The first was the explosion of data. From the early 2010s, every match in Europe's top leagues produced thousands of data points: every touch, every metre covered, every direction run. What in 2026, when I joined the sports desk of Belgrade Television, had to be counted by hand off videotape, is now computed before the referee blows the final whistle.
The second revolution gets less attention: the explosion of demand for conclusions. More matches, more platforms, more studio shows, more analysis channels. Every platform needs a take. Every take needs a person to deliver it. Every person needs a conclusion, every conclusion needs a reason, and every reason needs a metric standing behind it.
Those two revolutions run in parallel and produce a paradox I call blank-dossier pressure. Rising data volume leads people to assume data always exists. When it does not, they do not conclude that it does not. They conclude that whoever looked for it was not good enough.
I understand why. Looking at a data table is like looking at a battlefield map: the smallest detail is an arrow. When the map is full, people forget the map always contained unmarked territory. And the most dangerous arrow in the world is the one drawn over ground the cartographer never walked.
What fills the gaps
Human beings cannot tolerate a void. This is basic psychology, and football is one of the most perfect environments for it to express itself.
Three mechanisms fill the blank cells.
The first is availability. When there is no data on a player, people use their most recent impression of him. A miss in the 88th minute of the latest match becomes evidence for a whole season. A fine goal in a qualifier becomes evidence for a career. A football watcher's memory carries extremely skewed weights: it stores exceptional moments better than ordinary ones, even though in every probabilistic model it is the ordinary moments that decide where the season ends.
The second is survivorship. We remember successes, because failures leave no echo. How many young players were highly rated after six matches and then vanished into a third division somewhere? No one counts, so no one remembers. The ones who made it get retold as examples. The true success rate of that kind of decision is therefore inflated many times over in collective perception.
The third is the small sample. This is the most dangerous mechanism, because it wears the costume of science. A player scores seven goals in five games. A manager wins six of his first seven. A team keeps four consecutive clean sheets. All of it has numbers, all of it has data, all of it has charts. Only one thing is missing: statistical meaning.
I once received a scouting report in which a player was described through fourteen metrics, all computed over four matches. Four matches. A season is thirty-eight. Four matches is barely more than ten percent of the sample, and for high-variance metrics such as completed dribbles or tackles, ten percent of a sample tells you nothing except which four matches someone chose to count.
Four kinds of gap
After many years, I sort the blank cells in a dossier into four groups. Each has a different cause, a different fix, and a different way of going wrong.
The sampling gap
Easiest to spot, easiest to handle technically. Simply not enough matches. A breakout player in a small league, a national team with only four qualifiers, a manager twelve games into the job.
The correct response is to widen the observation window, not the interpretation. The widest gap in this trade runs between people who know they are looking at a small sample, and people who know it and conclude as though it were large.
When I built a database of twelve hundred attacking patterns covering the 2026 World Cup through the end of the 2026-20 season, I spent eight months working purely on rolling windows. For a metric such as ball recovery rate after losing possession, how many matches does it take to stabilise? The data says roughly twelve to fifteen. Which means any conclusion drawn from fewer than twelve matches sits inside the noise, no matter how many charts dress it up.
The category gap
Subtler, and far less acknowledged. Not missing data, but data collected in the wrong category.
For years, tracking systems counted passes, pass completion, touches very carefully. They did not count the thing that decides matches: the interval in which a player stands in the wrong position. A defender three metres off his line for four seconds can concede a goal, and no metric records those four seconds, because nobody touches the ball during them.
I call this the silent gap, because it exists in silence. The table looks complete. The reader feels satisfied. It is only that the most important thing is not in the table.
This is precisely the gap that drove me to build a geometric notation system of twenty-seven pressing patterns. I did not build it to count more. I built it because existing metrics cannot describe the geometry of a gap. The ball travelling through a gap does not mean the gap existed when the ball travelled; it existed three seconds earlier, when a midfielder decided to step up half a metre.
The access gap
Mundane but dominant across the industry: the analyst cannot see what needs seeing. No dressing room. No medical data. No internal reports. No training sessions.
This matters more than people think. Public physical data says a player covered eleven kilometres. It does not say he covered those eleven kilometres carrying a hamstring issue for six weeks. A table without a medical column tells the story of a different player from the actual one.
Before analysing Germany at the 2026 World Cup I had full positional and duel data. I had no data whatsoever on the true condition of the centre-backs after a long season in Germany and Spain. That was a gap I knew existed, and I flagged it clearly as unassessed. Many other analysts had the same data and flagged nothing.
The definitional gap
The hardest kind. Before asking whether a team is strong, you must define strong. Before asking whether a goalkeeper is good, you must define what a good goalkeeper does.
Most football arguments are not arguments about numbers. They are arguments about definitions, disguised as numbers. Two people look at the same table; one rates a keeper by save percentage, the other by goals conceded relative to expected goals against. They will reach opposite conclusions, and neither will be wrong on the arithmetic.
Germany 2026: when the table was right and still incomplete
On 27 June 2026, Germany lost 0-2 to South Korea in Kazan and went out in the group stage for the first time since 2026. Eighty years, once.
Before that match I published an analysis built on my notation system. My data showed the German defensive line holding an average position of sixty-two metres from its own goal in organised attacking phases, significantly higher than the safe band they had maintained at previous tournaments. Mats Hummels and Jerome Boateng were winning only about forty-eight percent of their duels across the first two matches. I wrote that a centre-back pairing winning fewer than half its direct contests, behind a line standing too high, was inviting exactly one thing: balls played long behind it.
The model was right. The article reached 870,000 reads.
Here is the part worth discussing. My analysis that day contained blank cells I had explicitly marked as unsupported by data. I had no metric for the quality of compactness in the second half of three matches played within eight days. I had no data on accumulated minutes per centre-back across the preceding season, because leagues publish minutes but not minutes under high fixture density. I had nothing to measure the fatigue of a squad that had won four years earlier and kept almost the entire structure intact.
I told the story of line height and duel rates. That story was true. But a system never collapses starting from the final defeat. It starts from decisions taken when nobody still feels pressure to change. Germany's 2026 World Cup began in July 2026, when a champion squad concluded it did not need to be different.
What I could tell was the geometry. What I could not tell was the time. And time is the part that explains why the same squad, the same manager and the same philosophy won four years earlier and were eliminated four years later by a team with no player at a European club.
Hakimi and the fourteen matches I watched twenty times
In December 2026, Morocco became the first African team to reach a World Cup semi-final. I tracked fourteen of their matches, including qualifiers and warm-ups around the tournament, rewatching some more than twenty times.
What I found was not in the table. Achraf Hakimi, a right-back, repeatedly left the wide channel and stepped inside when his team did not have the ball. The result was a defensive block whose shape opponents could not predict: the back four narrowed into three, and the midfield stretched into five with Hakimi in the right inside channel. For any opposing wide midfielder, identifying who was covering him became an equation with no fixed answer across ninety minutes.
Morocco conceded once in their first five matches, and that goal was an own goal. They eliminated Spain in the round of sixteen on penalties, with Hakimi himself taking the decisive one as a Panenka. They eliminated Portugal 1-0 in the quarter-final. They lost 0-2 to France in the semi-final on 14 December, then 1-2 to Croatia in the third-place match on 17 December.
My Morocco analysis video reached 1.2 million views on a platform in Beijing, and I sat as a studio guest on a national broadcaster throughout the tournament.
Now the blank cell. In my Morocco dossier I had positional data, interception counts, pass completion. I had no data at all on whether the pattern was repeatable. A team playing six matches at Morocco's intensity is not a large enough sample to conclude anything about a system. It is a phenomenon, and a phenomenon differs from a system in exactly one respect: systems repeat, phenomena do not.
My 2026 failure taught me more than every win that followed. That year I wrote my first piece for an emerging online sports platform, about the Chinese Super League, and got 312 reads and five comments. I did not quit. I went back through eighty Shanghai SIPG match recordings over three months and found something nobody was writing: the space between their midfield line and their defensive line was a fatal weakness, and across the 2026 season seven goals were conceded out of exactly that channel.
The lesson was not that I found seven goals. The lesson was that it took me three months to find seven goals, while it takes others three minutes to write that a defence lacks concentration.
The goalkeeper and the gap that got sanctified
There is a large blank cell in goalkeeper analysis, and it has existed for more than a decade.
In June 2026, Manchester City signed Ederson from Benfica for around 35 million pounds. In July 2026, Liverpool signed Alisson from Roma for 66.8 million pounds, then a world record for a goalkeeper. In August 2026, Chelsea triggered Kepa Arrizabalaga's release clause at Athletic Bilbao for roughly 71 million euros, breaking Alisson's record less than a month later.
Three transfers in fourteen months, three consecutive record fees. The stated reason was nearly identical in all three cases: distribution.
I do not deny the value of distribution. I deny the ratio between the value assigned to it and its actual effect on results. In my database, goals conceded directly from a goalkeeper's poor distribution make up a very small share of total goals conceded in top leagues. Goals scored from a goalkeeper's excellent distribution are similarly small. Goals conceded from an unsuccessful save make up the majority.
Which means the thing paid for most is not the thing producing most points. This is not a moral observation. It is a structural one: when a skill becomes easy to measure, it becomes easy to sell, and when it becomes easy to sell, it gets priced above its true value. Distribution is easier to measure than close-range reflex, because distribution is an intentional action, while reflex is a reaction dependent on whether ten other players were standing in the right places.
An average goalkeeper scouting template I once saw contained thirty-two metrics. Eighteen related to distribution, nine to organising the defensive line, five to shot-stopping. That 18-9-5 split was a decision made by whoever designed the template, and every reader of the template is influenced by it without knowing.
The error is not that eighteen metrics cover distribution. The error is that five metrics are treated as sufficient for shot-stopping. A top-league goalkeeper faces roughly one hundred to one hundred and thirty saves a season. One hundred and thirty samples. That is large enough for serious analysis, and it gets compressed into five rows.
Twelve hundred patterns, 23 percent, and the thirty-second window
In 2026 the major leagues stopped. Stadiums emptied, calendars scattered, press conferences moved onto screens. Through that period I did not watch a single match live.
I spent eight months building a database of twelve hundred attacking patterns, drawn from the 2026 World Cup through the end of the 2026-20 season. Every pattern was hand-coded against the same criteria, including the ones that failed. This matters: most public datasets only keep attacks that ended in a shot, because those are the easy ones to identify. Attacks that died in midfield are discarded, and those are the majority of football.
Then I tested it in Python.
The most important result is a number I still repeat years later: teams that pressed actively within the first thirty seconds of losing the ball recovered possession at a rate 23 percent higher than teams that pressed more slowly. Not 23 percent more pressing actions, but a 23 percent higher probability of recovery from the same starting situation.
That thirty-second window is the whole story. Before thirty seconds, the opponent has not organised. After thirty seconds, they have, and every recovery attempt becomes more expensive physically and cheaper in effect. Football does not reward the player who runs most. It rewards the player who runs inside the correct thirty seconds.
I do not believe in luck. I believe in 23 percent showing up a second time.
Yet inside that fifteen-page study there is a section I had to write in the language of emptiness. I lacked sufficient data on matches played after leagues resumed in the summer of 2026, because abnormal fixture density invalidated comparison with previous seasons. I lacked data on whether the 23 percent effect depends on fitness or on squad quality, because in my sample the fast-pressing teams were also the better-staffed ones. I had no way to separate those two variables with existing data.
That is the largest blank cell in my entire research career, and I left it blank. Anyone reading those fifteen pages sees the note on model limits. Most readers skip it. Most people who cite the study skip it too.
The mathematics of emptiness
There is a simple calculation I always present in training sessions for young journalists.
Assume a metric with a coefficient of variation around the football average, roughly 0.6. Across five matches, the standard error is so large that a ninety-five percent confidence interval spans almost the entire scale of that metric. In other words, after five matches you cannot distinguish a good player from an average one. Across thirty-eight matches the interval narrows enough for two players to differ statistically. That is the entire difference between a meaningful report and a meaningless report presented beautifully.
The problem is that big transfer decisions do not wait thirty-eight matches. The transfer market runs on a different clock from the statistical one. A player who shines across a four-week tournament triples in price within twenty days, and no metric prevents that, because the buyer is not buying thirty-eight matches of data. The buyer is buying a feeling about the last four weeks.
This is why a team dies before the match begins, at the negotiating table and on the transfer paperwork. No defeat starts in the first minute.
The cost of a wrongly filled cell
The cost of admitting you have no data is a silence in the article, an N/A line in the report, an answer the interviewer does not want to hear.
The cost of filling that cell wrongly is a thirty-million-euro contract based on six matches, a manager sacked after seven rounds because last season was judged on its final four months, a youth player promoted too early after three good games and gone by twenty-five.
In my career I contributed to at least three decisions I knew rested on insufficient data. I will not plead that I had no alternative. I did. I could have written that the sample was too small and accepted that the piece would be less attractive.
The hardest thing in this trade is not finding the truth. It is publishing something more uncomfortable than the truth: that there is not yet any truth.
The counter-intuitive angle: more data creates more gaps
The industry's natural response to missing data is to collect more. Invest in cameras, tracking systems, analysis departments, software. I have watched clubs spend large sums on datasets nobody inside the building can read to the end.
There is a paradox here that I have observed and never seen explained satisfactorily.
The more categories you measure, the more the unmeasured categories stand out, and the more scattered the reader becomes. A table with ten metrics makes people think about ten things. A table with three hundred makes them think about the first ten, then choose unconsciously according to display order. Display order becomes a strategic decision made by a software designer, not a coach.
Conversely, the clubs I have observed closely and judged to work best are not the ones with the most data. They are the ones that know precisely where they are blind. They keep a short list of questions their data cannot answer, and they assign humans to watch those specific questions live. They accept paying for managed blindness instead of paying for false certainty.
Sporting culture does not live in the stands; it lives in how people protect the shirt. And in the analysis room, protecting the shirt means leaving a blank cell blank when the blank is the truth.
The one thing I keep
Back to that night in Beijing with forty-seven empty cells.
I sent the report with nineteen cells filled, twenty-eight marked as unsupported, and a four-line preamble stating that any conclusion about this opponent should be treated as provisional. The recipient called back and asked if I was sure. I said the only thing I was sure of was the list of things I did not know.
Those twenty-eight blank cells did not weaken the report. They made it more accurate. The problem is that almost nobody wants to read a methodologically accurate report if it refuses to hand them a clear answer to act on.
Data does not lie, but it chooses who gets to hear it. In my trade, the people it chooses are usually the ones willing to sit still in front of a blank spreadsheet longer than everyone else.
This season is long. There will be more dossiers in which every category returns insufficient information. There will be more clubs deciding before the data arrives, and more matches explained by metrics attached to matches nobody watched.
What I want to know next round is not who wins. It is how many of the reports sent out this week left their blank cells intact, and how many had already been filled in by someone with a plausible-sounding story.
