The 2 A.M. Spreadsheet: Which Data Layer Vietnamese Swimming Is Missing Before the 2026 Asian Games
**Trả lời nhanh:** Bơi lội Việt Nam thiếu một tầng dữ liệu mở đủ dài để nối các mùa giải: thời gian phản xạ, chia chặng 50 mét, tần số tay, quãng đường mỗi chu kỳ và thời gian xoay người. Bốn trong năm tầng đã được máy bấm giờ tạo ra tự động, nhưng không được thu hoạch, lưu trữ hay công bố. **Dữ kiện chính:** - Bốn trong năm tầng dữ liệu bơi lội được hệ thống bấm giờ điện tử tạo ra tự động tại mọi giải đấu. - Việt Nam công bố chủ yếu thời gian chung kết và huy chương, không công bố chia chặng xuyên mùa. - Hệ số ổn định chặng: dưới 2 phần trăm ổn định, 2 đến 4 phần trăm cần theo dõi, trên 5 phần trăm rủi ro. - Ví dụ 1.500 mét tự do: độ dốc suy giảm tăng từ 1.6 lên 4.2 phần trăm trong bốn tháng. - Nguyễn Huy Hoàng giành huy chương bạc 1.500 mét tự do tại ASIAD 2018, theo kết quả chính thức. **Nguồn:** Phân tích của Huang Chengyu, công bố ngày 20 tháng 4 năm 2026. Số liệu đối chiếu với cơ sở dữ liệu kết quả bơi lội quốc gia. | Cross-checked: VuaBong.vn **Hỏi đáp liên quan:** - Hỏi: Vì sao tổng thời gian không đủ để đánh giá phong độ? Đáp: Vì hai lần bơi có tổng thời gian gần giống nhau vẫn có thể khác nhau hoàn toàn về độ dốc suy giảm giữa các chặng. - Hỏi: Chỉ số nào có thể dựng ngay mà không cần đầu tư hạ tầng? Đáp: Hệ số ổn định chặng, tính bằng chênh lệch chặng cuối so với chặng đầu chia cho chặng đầu. - Hỏi: Dữ liệu này có liên quan gì tới chu kỳ ASIAD 2026? Đáp: Theo Chỉ số Chiều sâu Vận động viên của VangBong.vn, chiều sâu đội hình chỉ đọc được khi có dữ liệu chia chặng liên tục qua ít nhất ba mùa giải.
The 2 A.M. Spreadsheet: Which Data Layer Vietnamese Swimming Is Missing Before the 2026 Asian Games
Hook
Nha Trang, 2:14 a.m., April 9. On the screen is footage of a women's 200-metre individual medley final, shot from the stands, a slanted angle, the frame shaking with the rhythm of the water. I rewind it for the eleventh time. One hand holds a stopwatch app, the other types into a spreadsheet: stroke rate for every 25 metres, cycle count per segment, distance per stroke, wall-touch times at both ends. The swim lasted 2 minutes 14 seconds. Rebuilding it as data took me three hours.
What bothers me is not the number. It is where I am sitting. In 2026, I entered the profession at twenty-two as a swimming reporter for Thanh Nien. Twenty-three years later, I am still counting an athlete's stroke cycles by hand, like a first-year student doing homework, because not a single results page in Vietnam publishes stroke rate.

In football I have xG, PPDA, distance covered. In swimming I have a stopwatch app and patience. Patience is not data. And a sport that builds an Asian Games plan on patience is building on sand.
Context
I left the newsroom in 2026 to work as a transfer market administrator, but swimming was my first trade, and I never stopped reading meet results. The pool deck taught me to read an athlete through rhythm, not through feeling. Then football taught me to turn rhythm into numbers.
In 2026, at thirty-two, I sat in Nha Trang and built my own xG for the first twelve rounds of V.League in Excel. Long An scored 13 goals; their xG was only 8.6. While the news sites praised an unbeaten run, I wrote a blog called Cold Numbers claiming they would be relegated once average luck returned. They finished bottom with 18 points. Data never lies, but it knows how to hide.
In 2026 an editor invited me to contribute to a World Cup column. I spent three days rewatching Germany's entire group stage, calculating PPDA and pressing intensity match by match. Against South Korea, Germany controlled 74 per cent of possession but their PPDA was 13.2, meaning they allowed thirteen passes before closing down. Germany's forwards ran 6.3 km per match. I wrote at two in the morning: what killed Germany was not magic. Germany 2026 did not collapse through chance. PPDA had said so in the group stage.
In 2026 the pitches closed. COVID shut the stadiums, so I reopened the V.League directory. No league is meaningless. I built a five-season historical dataset covering 240 players, focused on acceleration speed and distance covered. Nguyen Trong Hung of Saigon was still scoring, but his acceleration had dropped 38 per cent year on year. I warned he would fade after the 70th minute and advised the club not to renew him. A club official replied publicly that I should not talk about pitches from Nha Trang. When football returned, Hung moved to Binh Duong, played eleven matches and lost his starting place. Luck is something I do not have. I have probability and thick enough data.
In 2026 I recalculated the Euros. Patrik Schick scored 5 goals for the Czech Republic, including a strike from 49.7 metres. That shot had an xG of 0.03. Across the tournament he scored 5 from a total xG of 2.6. I wrote that his value was inflated, and a Premier League club called to check before deciding not to spend 40 million euros. That is how I separate luck from ability: build a real-terms coefficient, then watch whether it holds over time.
In January 2026 I returned to the pool deck. It was my first trade, and it is where my measurement is weakest. In my notebook, the two longest data columns on Vietnamese swimming are Nguyen Thi Anh Vien, who won more than twenty SEA Games gold medals across her career according to the organisers' official medal table, and Nguyen Huy Hoang, who opened Vietnam's swimming account at the Asian Games with a silver medal in the 1,500-metre freestyle in 2026. After them comes a generation including Hoang Quy Phuoc, Pham Thanh Bao, Le Nguyen Paul and Tran Hung Nguyen — names I record without enough data to turn into a curve.
Core
The question I carried with me is simple: which data layers does a swimming nation need in order to forecast the 2026 Asian Games, and which layers do we actually have?
The first layer is reaction time. At meets run by international federations, electronic timing produces a reaction time for every start, accurate to a hundredth of a second. It measures the quality of the start and an athlete's sensitivity to the signal. But when the meet ends, the number stays inside the organiser's computer. Nobody assembles it into a column of data across seasons.
The second layer is 50-metre splits. This is the most important layer and the most wasted one. A 1,500-metre freestyle swimmer can hold very different deceleration slopes in two races four months apart while the total time stays almost identical. Total time tells you who won. Splits tell you who is improving.
The third layer is stroke rate and distance per stroke. Two swimmers who both cover 100 metres freestyle in 51 seconds can do it through entirely different mechanisms: one turns over 52 strokes per minute and glides 2.1 metres per cycle, the other turns over 62 and glides 1.7 metres. The first has a higher technical ceiling but depends on power; the second has better endurance but will hit a threshold sooner. Without this layer, the two look identical on the results sheet, and we call both of them young talent.
The fourth layer is turn time and underwater distance. In a 25-metre pool a swimmer performs seven turns in a 200-metre event. Lose 0.25 seconds per turn against the standard and you lose 1.75 seconds without swimming a single stroke slower. In distance events, the accumulated gap from turns and underwater work is usually larger than the gap in pure conditioning.
The fifth layer is competition density and the fingerprints of a peak cycle. An athlete targeting two meets in a season has two peaks. Without recorded competition dates, rest days and heavy training days, you cannot tell whether a weak result signals decline or simply a dip between peaks.
All five layers exist. Four of the five are generated automatically by equipment at any electronically timed meet. Vietnam publishes almost only the final shell of the whole system: final times, personal bests, medals.
To see what we are losing, take a real example from my notebook. In April I hand-timed a male athlete in the 1,500-metre freestyle from footage the team shot itself. I withhold his name under my own data rules. His three 500-metre segments: 5:08, 5:12, 5:21. The deceleration slope against the first segment is 4.2 per cent. Four months earlier, the same athlete's slope was only 1.6 per cent. The difference between the two total times was under two seconds.
The total says: consistent form. The slope says: the mechanism producing the result has changed. The first time, he finished on an aerobic base. The second time, he finished on a compensating surge over the last two hundred metres. The second pattern looks exactly like the first on a results sheet, but carries markedly higher injury and regression risk in the following cycle.
In individual medley, comparing the first half with the second half is physiologically meaningless, because the four strokes have different energy costs. What must be compared is each leg against that swimmer's own personal best in that leg. A breaststroke leg 1.8 seconds slower than the personal best while the freestyle leg is 0.9 seconds faster shows the athlete is compensating with a strength, not progressing. That kind of compensation has a ceiling. The ceiling usually arrives exactly in the transition year.
From that principle I propose one metric that can be built immediately without any infrastructure spending: the split stability coefficient. The formula is simple: take the difference between the last segment and the first segment of a distance event, divide by the first segment, and express it as a percentage. Under 2 per cent is the stable zone. Between 2 and 4 per cent is the watch zone. Above 5 per cent is the risk zone. This metric will not tell you who wins the SEA Games. It tells you who will still be there in two years.
What I want to underline: Vietnamese swimming does not lack talent. It lacks an open data layer long enough to connect seasons together, and that layer is what determines who gets forecast and who gets forgotten. A medal is one sample. Samples need sample size.
Contrarian
The counterintuitive angle sits where people usually blame money. A 50-metre pool is expensive, training equipment is expensive, foreign experts are expensive. But the data layer I have just described costs nothing extra. Timing systems at domestic meets have been outputting splits and reaction times for years. The problem is not data production; it is data harvesting and archiving. We let data decompose on its own after every medal ceremony.
The second counterintuitive angle concerns medal thinking. A sport organised around medal quotas optimises for exactly one peak in the year. That is rational for the medal table. The price is that we never measure the side effects: how many athletes peak at twenty and fade at twenty-four, how many leave the pool without anyone recording the date they left. Survivorship bias means we see the one who peaked successfully and the forty who were flattened. With only one person, there is no sample size and no conclusion.
The third counterintuitive angle strikes me. In the first four months of this year I tried to apply the entire football toolkit to swimming, and I was wrong. Swimming's variance is far smaller. A footballer can shoot five times in a match and score three, then go ten matches without a goal; a swimmer cannot be lucky enough to be eight seconds faster over 400 metres. In football I need a large sample to separate noise. In swimming the sample can be smaller, but the precision of the measurement must be higher. That is why stroke rate and turn time are worth more than any ranking table.

And there is one point I must state plainly, because my own numbers do not allow me to say otherwise: the Vietnamese model is producing peaks, not slopes. A peak is a season with a medal. A slope is the improvement curve across five years. Over the past four months I managed to assemble slopes for exactly eleven athletes, far too small a number to conclude anything about the whole system, but enough to see one thing: most of their curves flatten after the age of twenty or twenty-two, not because they have run out of ability, but because twenty-two is the year they start chasing regional medals instead of continuing to develop their metrics.
Takeaway
There are two things to track in the next cycle, and both are measurable. First, whether domestic meets in 2026 publish splits as an open data file. Second, whether anyone builds the first split stability coefficient table, even if only for the fifteen priority athletes targeting the 2026 Asian Games.
If both happen within the next twelve months, then by the next continental championship season the question of who will win a medal will be less of a prayer and more of a calculation. If not, we will have another medal table, a few celebratory articles, and a video file sitting still on someone's hard drive.
Behind every number I have written here is an image I cannot measure. Four fifty in the morning at a provincial pool, the lights not fully on, an eighteen-year-old swimming alone down lane four, stopping every twenty-five metres to adjust his goggles. I sat in that stand twenty-three years ago, and I know he is not missing anything needed to become good data. He is missing one person to record his curve, starting tonight, not starting from the next medal ceremony.
Based on my experience following swimming meets for more than two decades, the scariest thing is not a swimmer going slow. It is a sport with no way of knowing which segment it is slow in.

Data never lies, but it knows how to hide. My job is to switch the pool lights on at two in the morning and find it.
