Trang chủSwimmingThe Blank Cells in Vietnamese Swimming's Data Sheet

The Blank Cells in Vietnamese Swimming's Data Sheet

**Core answer** Bơi lội Việt Nam công bố đầy đủ kết quả cuối cùng nhưng thiếu dữ liệu kỹ thuật như split 50m, thời gian quay đầu và nhịp quạt tay. Khoảng trắng này khiến giới phân tích không thể xác định nguyên nhân thành tích, cũng không thể dự báo phong độ ở các vòng đấu tiếp theo. **Key facts** - Bảng dữ liệu công khai của bơi lội Việt Nam có kết quả chung kết và thứ hạng nhưng thiếu split theo từng 50m. - Luật World Aquatics cho phép bơi dưới nước tối đa 15m sau xuất phát và sau quay đầu ở bơi tự do và bơi ngửa.

The Blank Cells in Vietnamese Swimming's Data Sheet

In June 2026 I sat in a rented apartment in Saigon, staring at Vietnam's three U20 World Cup group matches in South Korea. Across all three games the team generated 2.1 xG. The only goal, scored by Quang Hai from a free kick, carried an xG of 0.08. I did not sleep that night, but the lesson was clear: data only has value when it exists.

Seven years later I opened a different spreadsheet. This time it was swimming, a sport I have followed since 2026, when I wrote for Thanh Nien newspaper. The sheet had 42 columns. Thirty-one of them returned blank cells. Not zeroes. Blanks.

A zero is information. A blank is an absence. They are not in the same category, and confusing the two is causing us to misread the entire picture of Vietnamese swimming.

The Blank Cells in Vietnamese Swimming's Data Sheet

Context: a system built to hand out medals, not to understand causes

Vietnamese swimming presents an easily verified paradox: the top layer of its data is dense, and the bottom layer is empty.

If you want to know who won the men's 200m breaststroke at a recent SEA Games, you will find it in thirty seconds. If you want to know how much faster that swimmer went over the first 100m compared with the second, you will find nothing. If you want to know his reaction time off the blocks, you will find nothing. If you want to know how many metres he travelled underwater after the turn before surfacing, the answer is still blank.

This is the structure I call the single-layer results sheet: it records who won but not why they won, and therefore it allows no one to forecast who will win next time.

I have checked this many times in my work as a data consultant. Vietnam's public swimming record has four layers. Layer one is the final result: finish time, placing, athlete name, provincial unit. This layer is complete and public. Layer two is the 50m split: almost never published at national level, and where it exists it sits on foreign organisers' websites rather than in any domestic database. Layer three is detailed technical data: reaction time, underwater distance after the start and after each turn, turn time, stroke rate, distance per stroke. This layer does not exist in the public system. Layer four is training load and physical data: weekly volume, sensor readings, muscle-overload markers. This layer exists but sits inside each unit's internal files, unshared, unstandardised, and therefore incomparable.

For comparison: Singapore runs a national sports data system on a single standard; Thailand's sports authority keeps centralised athlete records; Malaysia's national sports institute collects periodic testing data. The three systems differ in quality but share one trait: they collect layer three and layer four data, not only layer one.

Vietnam has a great deal of layer one. And almost no layer three.

The technical layer: what nobody measures, and what it costs

An ordinary viewer watches a lane and sees one continuous motion. A data analyst watches the same lane and sees at least seven separately measurable events: the starting signal, the reaction off the block, the flight phase, the entry, the underwater kicking phase, the breakout, and the sequence of turns at every wall.

Each of those events can be broken into a number. Each number can be trained on its own.

In freestyle and backstroke, the rules allow a swimmer to travel up to 15 metres underwater from the start or the turn. This is one of the largest sources of separation between elite swimmers. An athlete who travels 13 metres underwater with a strong kick can claw back two tenths of a second against one who surfaces at 9 metres. Over 100m, two tenths of a second is the distance between gold and fourth place.

I know this. But I cannot prove it with Vietnamese data, because that data does not exist. And this is the crucial point: a conclusion that cannot be verified is not a weak conclusion; it is simply not a conclusion.

The same applies to turn times. In a 25m pool, a 200m race contains seven turns. If each turn is 0.15 seconds slower than a rival's, the total loss exceeds one second, enough to completely reorder the standings. Yet nobody in Vietnam publishes average turn times. Nobody publishes stroke rate. Nobody publishes distance per stroke.

Stroke rate and distance per stroke are structural variables. Their product produces speed. A swimmer can accelerate by raising stroke rate, by increasing distance per stroke, or by both. Those two paths carry completely different physiological consequences: higher stroke rate burns energy faster and suits short events; greater distance per stroke demands a better technical base and suits long events. Without those two numbers, there is no way to know whether an athlete is on the right path or simply burning through reserves needlessly.

In the dataset I opened, all seven technical variables above returned blank. The only technical conclusion available to me is: insufficient information to assess.

The performance layer: times exist, frames of reference do not

This is the second paradox, and it is subtler.

Times exist. You know how many seconds a Vietnamese swimmer took over 200m freestyle. You know the placing. But a lone time figure means nothing without three things: a reference frame, pool context, and sample stability.

The reference frame begins with records: world, continental, national. Vietnam has recognised national records, but the speed of updating and cross-verifying them across sources is very slow. I have encountered three different versions of the same national record at three different sources, separated by a few tenths of a second, with none of them stating the date it was set.

Pool context matters more. The same swimmer over the same distance in a 50m and a 25m pool produces two different results. A 25m pool contains more turns, so a swimmer with strong turning technique gains an advantage; a 50m pool contains fewer turns, so endurance and long-course technique dominate. A good short-course result does not automatically convert into a good long-course result. Without knowing the pool length, half the value of the time figure evaporates.

The third factor is the equipment era. In 2026 World Aquatics banned high-tech polyurethane racing suits. Records set before and after 2026 do not share the same comparative value. A 2026 record set in a high-tech suit sitting beside a 2026 record set in a textile suit are two different quantities. If the sheet does not record the year and the suit type, it is blending two eras into one column.

Finally there is pacing structure, and this is where Vietnamese swimming is thinnest, because it depends entirely on splits. A swimmer who goes out faster than they come home is on a positive split; one who comes home faster is on a negative split. A positive split usually signals an athlete who went out too hard and paid at the end; a negative split signals rhythm control and a good endurance base. The two demand entirely different physical programmes. Without splits, coaches are guessing.

Competition structure and entry pathways: reading results without knowing the tier

A common domestic error is to place results from different tiers side by side without discounting them.

The Olympics, long-course world championships, short-course world championships, World Cup, continental championships and national championships each carry a different weight. A result at a national meet in a year without a major peak is often swum through, meaning the athlete was training rather than racing for a result. That kind of result should not be read as peak form.

Conversely, a result at an Olympic qualifying meet carries far higher value, because the entry standards are set by World Aquatics and the psychological pressure coefficient is large.

Vietnam's selection mechanism rests on international time standards. In principle an athlete meeting the A standard enters directly; one meeting the B standard depends on quota allocation. What interests me more is the internal selection mechanism: who decides which athlete goes abroad for training camps, on what criteria, and which data is used to make that call. With a database that only has layer one, camp decisions are almost certainly made on impressions and relationships rather than evidence.

I am not accusing anyone. But the logic is plain: if you do not hold layer three and layer four data, you cannot make decisions based on layer three and layer four.

The regional map and the talent supply chain

In Southeast Asia, swimming has held a fairly stable order across many SEA Games: Singapore in the lead group, then Thailand, Malaysia, Indonesia and the Philippines, with Vietnam in the chasing group. At continental level, Japan, China and South Korea dominate the Asian Games swimming medal table almost entirely.

What stands out is that the gap between the SEA Games group and the continental group is far larger than the gaps inside the SEA Games group. That is a signal about the structure of the talent supply chain, not about individual effort.

Vietnam's talent pipeline runs from provincial gifted-sport schools up to the national youth team and then the senior national team. The depth of the system depends on how many children reach a pool between the ages of six and ten. That is an infrastructure variable, not a talent variable.

In many provinces, four-lane pools are built as public works for drowning-prevention swimming education. That goal is entirely correct and necessary. The side effect is that the supply of elite athletes is capped by facilities, because there is no competition-standard pool for long-course training. An athlete who wants to train with an internationally qualified coach usually has to relocate to one of a handful of large centres.

Alongside that, recent personnel movement carries a signal: large centres hire foreign specialists in some physical and technical roles, while domestic coaches mature slowly. That is an observation at the human level. It is also a reason data does not accumulate: whenever a specialist leaves, knowledge of the athletes leaves with them unless an archiving system exists.

Rules, officiating and anti-doping: silence is not emptiness

Three rule groups affect swimming results directly, and Vietnamese data rarely records them.

The first is the 15-metre rule. Travelling beyond 15 metres underwater after the start or a turn is a foul. It is a permanent risk point in freestyle and backstroke, especially when an athlete pushes the underwater kick advantage to the limit.

The second is breaststroke technique. Breaststroke has strict rules on the kick cycle, the sequence of movements, and the number of underwater kicks after the start. An otherwise optimised breaststroker can be disqualified over a small technical detail, with no data sheet retained to learn from it.

The third covers the backstroke starting device and competition equipment rules.

On anti-doping, governance involves World Aquatics, the World Anti-Doping Agency and the national anti-doping body. One principle I hold absolutely in every analysis: silence in the data is neither evidence of a violation nor evidence of compliance. Neither can be inferred from the other. Without a confirming document there is no conclusion, including a favourable one.

The career curve: the puberty barrier and a blank that gets misread

This is the section I want to dwell on most, because it bears directly on the interests of young athletes.

In women's swimming, the puberty barrier is the single most important screening factor between the ages of 13 and 17. Changes in body composition, fat distribution, height and arm span can stall an age-group champion for two or three seasons. It is a universal physiological phenomenon in every country with a youth swimming system.

In Vietnam this phenomenon is observed constantly but almost never recorded numerically. Here is the consequence: when a 15-year-old female swimmer fails to improve over a season, the system holds no data to distinguish between three entirely different possibilities. First, she is passing through a normal physiological barrier and will return. Second, the training programme no longer fits her new body. Third, there is an undetected injury or psychological problem.

Those three possibilities require three different responses. Without layer four data, all three collapse into one label: no progress. In football I once watched an impression-based judgement damage a young player's career. In swimming the damage can be worse, because an athlete's career window is shorter than a footballer's. The career span of an esports professional is shorter still, and post-retirement support structures there are close to zero. Three sports, one shared data failure.

On injury, the two most common occupational syndromes are shoulder injury from repeated rotational loading in freestyle, butterfly and backstroke, and medial knee injury in breaststroke from the kicking action. Both have early warning markers detectable through training-load data.

In 2026, when pools and stadiums closed because of the pandemic, I took part in reviewing movement data for a group of players in Saigon. One finding stood out: high-speed running distance rose by roughly twenty per cent in the period before a muscle injury occurred. We used that signal to split training into four load thresholds, and the following season recorded a clear drop in muscle injury cases.

Let me be explicit: we did not conclude that running fast causes injury. The more plausible mechanism is that once the body has accumulated fatigue, the athlete still tries to maintain volume, compensatory movement patterns appear, and those patterns load one muscle group locally. The higher speed reading is a surface marker, not a cause. This is the principle I apply to every data series: before saying X leads to Y, you must identify the physical or behavioural mechanism linking the two variables. If you cannot, you may only say they are associated.

The risk profile: and one risk that sits outside the pool

When I build a risk table for Vietnamese swimming, the professional risk cells stay empty for lack of input data. Competitive risk cannot be assessed. Rules risk cannot be assessed. Doping risk cannot be assessed, and I refuse to speculate in that cell at all. Psychological risk cannot be assessed.

Only one cell can be filled, and it sits outside the pool: analysis-integrity risk. When a dataset returns blanks in most columns while keeping its formatting structure intact, it creates an illusion. A reader skimming it sees complete section headings and tidy tables, and readily assumes that the empty places were checked and found unproblematic.

In data governance that is the most dangerous class of error: the silent failure. A pipeline that raises an error gets fixed. A pipeline that returns empty results in a format-valid structure gets used.

Public narrative and the expectation gap

Vietnamese sports media has a repeating pattern: after every SEA Games, a young athlete with a good result gets labelled the successor to a major name.

The label is not entirely harmless. It creates a measurable expectation on the public and sponsor side, while the athlete holds no data at all with which to confirm or refute that expectation.

The structure sits here. In Vietnamese swimming a handful of names have generated enormous expectation headroom over more than a decade. One female swimmer competed at an Olympics at sixteen and won a substantial haul of SEA Games golds across several consecutive editions; domestic media have credited her with more than twenty golds at that event. One male swimmer earned outright Olympic qualification in two long-distance freestyle events. Another male swimmer holds national records in sprint events. A younger swimmer has won SEA Games medals in the individual medley.

Those are real milestones, and they set an extremely high benchmark. The problem is this: when a 16-year-old swims a good time, the data system cannot tell us where that athlete sits on the development curve. Without layer three and layer four data, we can only compare final times. And comparing final times at 16 between two athletes at different points of puberty is biologically invalid.

I have heard many such judgements. Every time, the same question returns to me: if the successor label is attached on the basis of one figure, what proportion of such labels have historically come true? I do not have the answer. And the fact that I do not have the answer is itself part of the problem.

Industry ripple: empty data propagates along the chain

Along the value chain swimming has three segments.

The upstream segment is child development, the swimming-lesson market and talent supply. It runs on how many children get access to water and on the quality of the curriculum. Data here barely exists publicly.

The midstream segment is athletes and competitions. This segment holds the most complete layer-one data, and it is where the absence of layer three does the most damage.

The downstream segment is media, sponsorship, equipment and derivative markets. Star effect here depends directly on the midstream: an athlete with a clear data story carries far more media value than one with only a finishing time, because a data story produces a curve over time, and a single figure does not.

In equipment, global brands operate to their own data standards and typically collect athlete data through sponsorship contracts. That means part of the layer three and layer four data already exists, held by sponsors rather than by federations or domestic coaches.

On the betting segment I offer no assessment whatsoever, and will not under any circumstances.

The contrarian angle: blank space is not neutral

Here I want to step away from conventional analysis and say something directly.

Blank space in data is not a neutral state. It is always read in some direction. And the default human direction of reading is the unfavourable one.

If we hold no data on a 15-year-old female swimmer who has stalled, the default reading is that she lacks potential. If we hold no data on a male swimmer's turn times, the default reading is that his technique has nothing worth discussing. Both readings are inferences from silence, and both can be wrong.

This is where I have to correct myself. I have a habit of treating every context as a variable that can be switched on or off. In football I modelled home advantage as a parameter and showed that when the crowd is neutralised, that advantage collapses to near zero. When the stands fall silent, home advantage dissolves into a number close to zero. But there is a share of variance my model cannot explain, and I must concede it: when crowd intensity exceeds historical thresholds, the number loses part of its force. That is a confidence interval I have to state openly rather than pretend the model covers everything.

Applied to swimming, this means: an athlete walking into a major final before thousands of spectators indoors, with sound bouncing off the pool hall ceiling onto the water, is in a state no spreadsheet can simulate. Some swim better in it. Some swim worse. We have no data to know in advance. And our lack of data does not mean the effect is zero.

Correlation also does not run backwards into causation. The fact that a country has more pools and more medals does not by itself prove that building pools produces medals. The intermediate mechanism must be identified: pools produce more children who can swim, larger numbers raise the probability of spotting talent, and higher probability produces more elite athletes over a long horizon. Remove any link in that chain and the conclusion collapses.

The point lies elsewhere

I have spent most of this piece on the columns that returned blank. What I actually want to say does not lie in those columns.

It lies here: Vietnamese swimming has produced athletes capable of competing at continental and world level for many years. Those results are real. But we hold them the way one holds a photograph, not the way one holds a diary. A photograph tells you something happened. A diary tells you how it was built, and what to do next time to build it again.

The shot appears once. Its trajectory lasts for years. In swimming, each start also appears only once in a round. But the trajectory behind it, the programme, the training load, the technique, the psychology and the schedule, runs for four to eight years. If we record only the moment of surfacing and ignore everything submerged beneath it, we are recording a very small part of the truth.

Ordinary observers watch goals to understand matches. I watch matches to understand years. In swimming, ordinary observers watch lanes to understand results. I want to read the data sheet to understand a whole generation. An era of technique dies when nobody reads its data sheet any more.

Signals for the next cycle

If you manage data at a swimming unit, the first task is not buying more equipment. The first task is placing a validation gate at the head of the pipeline: any dataset returning without splits, without turn times, or without at least one technical variable must be flagged as not yet eligible for analysis. Not free of problems. Not yet eligible.

If you are a coach, the cheapest and most effective action is recording splits manually with a stopwatch at every test session. Two columns of data, recorded consistently over a year, are worth more than any analytics software.

If you are a fan, be wary of tidy tables that are mostly empty cells. A fully structured sheet does not guarantee fully populated information.

Every shock has its own probability. We call it a shock when we have not yet checked the table. But when the table returns blank, what we call a shock may simply be something we were never patient enough to measure.

And this is what I believe: over the next decade, whichever country records layer three and layer four swimming data consistently will no longer have to talk about the class gap in the language of emotion. They will talk about it in curves. Football and esports differ not in essence but in reflex speed, and both are waiting for a generation that knows how to record itself.

Cầu thủ liên quan