Yihua Zhang
Articles
Posts
Today I finally cleared up a small misunderstanding I had been carrying around in my grasp of music theory. The question that led to it was whether the sixth in a Cm6 chord is A or Ab. Plenty of people will react straight away and say that it is obviously A, the note a major sixth above the root, so how could it possibly be Ab. But my understanding of the Cm family of chords had always run on a different logic, and it was only after discussing it with ChatGPT today that I found my understanding was wrong. I had always believed that as long as a chord is written CmX, whatever X happens to be, the extensions stacked in thirds should all be taken from the C minor scale. That reading was in fact quite self-consistent for a while. Take Cm7, which is C, Eb, G, Bb. No problem there, because Eb and Bb both belong to the key of C minor. This logic even looks fine all the way up to Cm11, because the second degree (the ninth) and the fourth degree (the eleventh) of C minor are D and F, exactly the same as they are in C major. The trouble shows up at Cm13, or at the simpler Cm6, because the sixth degree is where major and minor finally part ways. In C major it is A, and in C natural minor it is Ab. It took a correction from Teacher GPT before I learned that the m in Cm governs only the third and has no coupling effect on the notes above it. The 2, 4 and 6 above it are all taken from the natural major scale, which makes them a major second, a perfect fourth and a major sixth. Quite an interesting mistake to have made. Having cleared it up, though, I feel that my understanding of the chord symbol system is more self-consistent than it was before.
Translated from Chinese by AI
Read in full →Today was the first time I had a chance to sit down and watch the trailer for HBO's Harry Potter series properly, and it went a long way beyond what I expected.
As a Ravenclaw who has read the Harry Potter books at least five times, I have always held some regrets about the films. Before I talk about the new series, I want to sum up what I think of the films.
First, taken as a whole, the Harry Potter films were undoubtedly a success. They pushed the Harry Potter IP to a new level of fame. Their soundtrack is also one of my favourite scores in film and television, and I would put it alongside Attack on Titan and Game of Thrones. As the first attempt to bring Harry Potter to the screen, the films reached a fairly high standard in building and presenting a visual world and in shaping the characters, and their portrayal of the various kinds of magic is satisfying enough, especially when you remember that this series of films is already more than ten and in some cases more than twenty years old. We should not take today's standards and denounce a work made ten or twenty years ago on every count. It is interesting that Harry Potter and the Philosopher's Stone (2001) and Yongzheng Dynasty (1999) are almost contemporaries.
Even so, I still want to talk about my regrets over the films.
First, we should be clear about one thing. Saying that the Harry Potter films are a success is a judgement made from the point of view that the films and the books are two different works. But once you take into account the faithfulness to the books that many book readers are after, the films have to be marked down a great deal.
That said, I personally think that reasonable changes to the source material, in the plot or in how the characters are drawn, are beyond reproach, as long as a complete story gets told in the end. But on that one point, the complete story, the Harry Potter films absolutely cannot be given a pass.
From the first film to the last, the Harry Potter films cut down heavily on parts of the books that are decisive for where the story goes, so many episodes that are both wonderful and important never made it onto the screen. Note that what is being judged here is not whether the books were reproduced in full, but whether the story itself was told completely.
The most jaw dropping cut of all is surely in Harry Potter and the Half-Blood Prince. The hidden thread in the book about Voldemort's origins, his rise and where the Horcruxes came from was almost entirely removed from the film. For example, how the young Voldemort found out about his own family, how he found the Gaunts, how he killed his own Muggle father, and how he came to choose those objects of special significance as Horcruxes. Together these memories explain why Voldemort made those Horcruxes, and why Dumbledore was able to work out step by step what they might be. This material is essential for understanding Voldemort as a character and for understanding how the story develops later on. Even allowing for the running time of a film, cutting plot of this importance on that scale still weakens the completeness of the story badly.
Harry Potter and the Goblet of Fire is another example. The film does explain at the end that Barty Crouch Junior had been impersonating Mad-Eye Moody with Polyjuice Potion, but how he got out of Azkaban, how his father kept him under control for years, how he briefly broke free of that control at the Quidditch World Cup, what part Winky played in all of it, and the complicated relationship between him and his father Barty Crouch, were almost all removed. Even the reasons and the course of events behind his killing his own father at the end were never really laid out. As for Winky the house-elf as a character, the Society for the Promotion of Elfish Welfare that Hermione founded to fight for house-elf rights, and Dobby's large part in this stretch of the story, they were either cut or handed over to other characters. That is in fact a very important part of what shows Hermione's character, her values and her growth. Rita Skeeter appears only a few times, in something close to a walk-on part. Her unregistered Animagus form, how she managed to get first hand news from places she could not possibly have reached, and how Hermione finally saw through her secret step by step, this whole line of the story was of course skipped as well.
The same problem is in fact even more obvious in Harry Potter and the Prisoner of Azkaban. The film keeps a prop as important as the Marauder's Map, and yet it removes almost the whole story of the four Marauders behind it. A viewer who has only seen the films can hardly understand why Lupin knows what the map is the moment he sees it, why the four people who made it are called Moony, Wormtail, Padfoot and Prongs, or why James could turn into a stag, Sirius into a dog and Peter into a rat. The reason they took the risk of learning to become illegal Animagi in the first place was precisely so that they could keep Lupin company when he turned into a werewolf. And that friendship from their school years directly explains the Shrieking Shack, the Whomping Willow, the Marauder's Map and even the Patronus later on. The films keep the results and cut away much of the cause and effect behind those results. For a viewer who has not read the books, these things often stay at the level of "this is some object in the wizarding world", with no way to understand why it is there at all.
You can tell even from the thickness of the books that J.K. Rowling's grip on the details of the whole work, and her design of foreshadowing and callbacks between the earlier and later volumes, are astonishing. Those details never made it into the films in full. So the regret I am describing is not a belief that the films should have matched the books exactly. It is that a viewer who has only seen the films may not have appreciated even half of how good the books really are. On the other points, such as the films' failure to portray Harry's growth as a person, the leads "usurping" other characters' parts, and the acting of Daniel, who played Harry, plenty of bilibili creators have already taken all of it apart in great detail, so I will not go into it here.
Back to the trailer for the series. From the trailer as it stands, there are pleasant surprises and there are frights. Most of it, of course, is the pleasant kind.
The pleasant surprises come first of all from the way it shows the details. Note that I am still not using the word "reproduction" here, because there are places in the trailer that do not match the books exactly. But at the very least you can see that the director really did think about these details, and thought about what they do for moving the plot along and for building the characters.
Beyond that, the acting the various characters are showing so far, the casting that sits closer to the descriptions in the text, for example that Hermione was never the prettiest girl in the story anyway, and the fact that characters who counted as minor figures in the films get a chance to appear again, all leave me looking forward to it a great deal. The fright, such as it is, comes mainly from the casting of a Black actor as Snape. Judging purely by the command of the lines that the trailer shows, this actor's ability is more than enough to carry the character of Snape. The problem is that Snape was one of the Death Eaters, and the Death Eaters are a group built on prejudice about blood purity. Given that setup for the character, casting a Black actor as Snape does produce a certain sense of dissonance, both in the cultural setting the character sits in and in how it looks on screen. That is probably why many viewers find the Snape casting hard to accept, even though this actor's classically trained acting is no worse than the Snape of the films. I find it hard even to picture the scene where Snape covers Lupin's class and teaches about werewolves with slides, a Black actor in black teaching robes, standing at the end in front of an entirely black background. As for the casting of the other characters, at this point I even think it is in no way inferior to the films. For some of the leads, there is the potential to go beyond the versions the films gave us.
In any case, remaking a work that has already been this successful and that has such a huge audience behind it must come with enormous pressure. The key question for this adaptation is how to make something fresh while letting the series build a character of its own. Fortunately, compared with the films, the series has at least two advantages that cannot be matched.
First, the series has more running time.
That means it can show more of the plot and more of what the books designed, and it can give more room to those characters who have vivid personalities in the books but were badly compressed in the films. Neville Longbottom's background and the story of his parents, for instance, or the story between Luna and Harry, or characters like Cho Chang who never got much room in the films, and a good number of the villains. All of this has a chance to be opened up again, which would make up for one of the biggest regrets about the films.
Second, unlike the period when the films were made, Harry Potter has long been finished.
That means it is now completely clear which details are foreshadowing, which places will echo something later, and which apparently unimportant pieces of the setting will matter in the end. I know that during the making of the films the screenwriters were told in advance about some later Harry Potter plot that had not been published yet, but even so, they could not possibly have had the complete overall view we have today of what foreshadowing exists inside the whole set of books and how it all finally resolves. For example, I remember that in the first novel, when Harry first sees the statues in the castle, the author describes them in considerable detail. Harry's first feeling at the time was that they might really be able to move. And in the Battle of Hogwarts in the last book, Professor McGonagall really does use magic to set all those statues moving, so that they join the fight and defend the school.
This is one of the most fascinating things about the Harry Potter books. Many of the details you take on a first reading as being there only to fill out the world turn out, several books later, never to have been details written down at random.
Translated from Chinese by AI
Read in full →Keigo Higashino's Newcomer is a very interesting novel. Seen from the reader's side, this classic puzzle mystery opens up along a completely different dimension, and it shows the whole process of a detective working a case, completely and truthfully. I think there are two main things that separate Newcomer from other mystery novels. One is the way the plot is moved forward, and the other is how true to life the account of the case is.
I think the process of reading the novel and gradually uncovering the truth of the case can be compared with looking at a painting. This time the author guides the viewer through the painting starting from the layers. Just as when you use Photoshop, we can bring in the idea of a "layer". Inside every layer, the author has to paint a picture that is complete and true in itself. That picture has its empty spaces and its solid forms, it has a main line and it has details, and yet it seems to have no direct connection with the theme of the painting as a whole. But when these layers are stacked together, you find that they are woven into one another, and in the end they make up one complete and truthful theme. Each time the reader starts again, he is not following one complete painting from beginning to end. He is looking at it one layer at a time. Every layer has a character of its own, and the next layer will often use some of the elements from the layer before it. Only when the last layer appears does the theme of the whole painting truly leap off the page.
As for what I mean by the account of the case being true to life, in mystery novels of the past the account of a case tends to stay on the key details that are directly related to the case, or that can push the case along. But when a detective actually works a case, he often meets statements that have nothing to do with the case itself and that also do not match the truth. Each of these false statements has a purpose of its own, and that purpose need not have anything to do with the truth of the case. These statements, true and false mixed together, have nothing to do with the truth of the case and yet really do affect the judgement of the person investigating it, and the detective has to pick them apart strand by strand to look for the real answer. Uncovering the reasons behind those false statements, and the process of uncovering them, is itself what shapes one real character, one real family and one real story after another.
Translated from Chinese by AI
Read in full →Who says there is no such thing as Mom's favorite child? I run two Claude sessions with exactly the same model and the same thinking effort, and I still faintly feel that one session is smarter than the other, and the important jobs all tend to go to the smarter one. LoL
Translated from Chinese by AI
Listening seems never to have been treated as a proper subject and trained seriously during our school years. I only really took that lesson in as an adult, while I was preparing for a German exam.
Many people may find this extreme claim odd. As a human being, from the moment I was born until now, the language I have heard must add up to more than the language I have spoken, so how can you say that I do not understand what I hear? At first I had never seriously thought about it either. Let us start from one phenomenon. For a typical Chinese student who wants to take the TOEFL (the standard English test that almost every foreigner has to take before coming to the United States, at roughly university lecture level), the two hardest parts to prepare are probably listening and speaking, and that much is common ground. But we usually think our listening is weak because we cannot remember the words, or because we are not familiar with the sentence structures, and we easily overlook one important question. If you were given information of the same complexity directly in your native language, and you got to hear it once, would you answer the same questions correctly?
That extreme claim actually comes out of my own experience of learning foreign languages. I already scored above 110 on the TOEFL in high school. I started learning German at university, and after four years of it I took the TestDaF and passed with 4x4, which is roughly the German equivalent of a TOEFL score of 100. Through the whole preparation, the strongest feeling I had was that when information came in, my brain often could not keep up with processing it.
This failure to keep up was not a failure to translate fast enough. It was that once information carries a certain complexity, and especially when you have to enter a context quickly through hearing alone, for example when a lecture on some topic suddenly starts and you have no idea beforehand what it is going to be about, I found it hard to keep processing the logical relations between one piece of information and the next. The previous sentence may not be fully sorted out yet when the next one has already arrived. Once one logical link is missed, everything after it tends to break off along with it. This weakness comes partly from a weak ability to process complex information in the first place, and partly, perhaps, from an instinctive escape reaction in the brain when it faces complex information, which is what we commonly call "mind wandering".
"Mind wandering" is a rather vague and inaccurate way to put it. Go back ten years. A secondary school student whose mind wanders in class seemed like the most normal thing in the world. When a teacher saw a student's mind wandering, the natural reading was that the student was thinking about something else, wanted to go out and play, or simply had no interest in the lesson. Looking back now, I think there may be another main cause that has never been noticed, which is that his ability to process information through hearing is too weak, so the mind wandering is itself an instinctive escape by the brain. In other words, it is not that the student has something else he wants to do. It may just be that his brain failed to keep up somewhere, and so it simply escaped.
Once this escape becomes a habit, the effect later on is that in similar situations the brain goes on escaping, beyond your control. In a listening exam, when the information gets complicated, your attention cannot hold. Even at a lecture you chose to attend because the subject interests you, once the content gets a little more complex, you cannot help drifting off. In a company meeting, when a coworker talks about something you are not familiar with, you slip away without noticing. Even later, when you are arguing or quarrelling with your partner, you may fail to follow the other person's logic and information correctly, so the quarrel moves from one topic to another and then goes on without end.
One point here seems important to me. These things very often happen "subconsciously", and even "against a person's own will". You may genuinely want to follow this lecture, genuinely want to take part in this meeting properly, and genuinely want to understand what the other person is actually saying, but once the information passes a certain level of complexity, the brain still starts to escape on its own. I feel more and more that these problems, which look completely different from one another, may all trace back to the same thing, which is that listening is too weak.
Go back to more elementary classrooms, primary school for instance. Every country has classes in its own native language, such as the Chinese language class for Chinese people, or the English class for Americans. They all offer language training that covers writing, reading and even speaking, and a listening class in any real sense is the one thing that hardly exists. The listening class I mean here is not students sitting there listening to the teacher talk. It is training specifically in how to process complex information through hearing, including analysing information in a complex context, reading emotion, following logical relations, and training attention. You can of course say that every class we take in our native language is a listening class, because students sit there listening to the teacher every day. But what those listening classes provide is not training whose objective is listening.
Anyone who trains large models knows that in multi-task learning, a gain on one task very often brings a drop on another. Training many tasks together does not mean that every ability will naturally improve, and what ends up being strengthened still depends on what your training objective is. By analogy, we are of course listening every day, but the direct goal of these classes was never to train how to listen. And if an ability has never been trained as an objective of its own, a large amount of repetition will not necessarily make it better. It may even work the other way round. What our classes that do not aim at listening may be doing is repeating, many many times over, the habits we already had when we take in information from other people. Good habits are reinforced, and bad habits are reinforced just the same.
So how should listening actually be practised? Personally I think podcasts are a good route. And practising listening in your native language may work better than practising it in a foreign language, because that removes as much of the difficulty of the language itself as possible and keeps the training on the act of listening. Compared with sitting through university lectures, podcasts cover a wider range of topics, many of them run long enough, and they can hold a certain depth. I think that when we train listening, what we want is to run into "skirmishes" of information fairly often. If you grew up playing Red Alert, the idea of a skirmish will be familiar. You do not know in advance what you are going to meet, and you can only take things in and judge them as the information keeps appearing. Listening training can be a similar process. What we want is that, with no preparation for the content at all, a deep and well layered talk suddenly starts, and then, once the podcast ends, we retell the layers of its logic as fully as we can. It is worth stressing that the small details do not really matter. Failing to remember one particular name, number or example does not affect the training. What you really have to catch is the logic of the whole podcast, or in other words, how this "story" is told step by step, meaning where it starts, what turns it takes along the way, why it moves from one question to another, and how it reaches its conclusion at the end. I think what this training really exercises is, on one side, the ability to hold attention for a long time, and on the other side, the brain's ability to keep processing and organising the information that hearing brings in and to build the logical relations inside it. In other words, what we are training may not be "how much you heard", but whether the brain can keep working and stay active for a long stretch while unfamiliar and complex information keeps coming into your ears, and take that information in properly.
Translated from Chinese by AI
Read in full →I reread Keigo Higashino's After School, and seven or eight years have passed since the first time I read it. I still cannot understand the murderer's motive. Teenage girls really are a mystery. Lately I have been reading Higashino's classic works, including The Murder in Mansion Masquerade, Naoko and The Home Where I Died, all of which are new to me. On bilibili, driving to work, I have also finished listening to Journey Under the Midnight Sun, The Devotion of Suspect X, The Red Finger and I Killed Him. The first two were rereads and the last two were a first listen. Listening to a book is a strange experience, and the strange part is that you know how a name is pronounced but you have no idea which two characters it is written with. The first time you hear a character's name, you may quietly give that name a spelling in your head, but as the case unfolds and you come to know the character better, you may start to feel that "these other two characters would actually suit this person's name better", so the character's name keeps changing as the story goes on. Then, once you have really finished listening and you open the book and look, you find that you seem to know none of these people, and when you look more closely it turns out that the author chose completely different spellings for the names that sound the same. Right now I am rereading Mysterious Night and Silent Parade. That is right, reading on two threads, LoL. Basically I keep two Apple Read windows open, and whichever one I happen to open is the one I read. Of all the works above, the one with the strongest aftertaste is still Naoko. Even though I could guess the ending halfway through, when I actually reached the last section I still could not help letting out a deep sigh. You find yourself hoping "please not this, please not this", but if it really had not happened, you would then be torn over "if not this, then what should it be?"
Translated from Chinese by AI
Read in full →About two months ago, an article called Don’t Paste the AI started circulating everywhere, both inside and outside the company. Its argument was simple: when a coworker asks you a question, you shouldn’t just paste an answer generated by an AI agent. A short, thoughtful response written by a human is almost always more useful. After all, if I wanted an AI answer, why would I ask you? I could just ask AI myself. The link spread almost overnight. Everyone seemed eager to share it, as if some long-simmering frustration had finally found the perfect slogan. Suddenly, everyone was a victim of coworkers dumping walls of AI-generated text into conversations. I didn’t share it. Not because I didn’t have an opinion, but because I disagreed. At the time, though, I didn’t say so publicly. Partly because doing so would have sounded suspiciously self-defending, as if I were defending my own reputation as someone who spends all day mindlessly forwarding AI output to coworkers. And also partly because I didn’t feel like being the person who walked into a room while everyone was enthusiastically agreeing with one another and announced that they were all missing the point. But two months have passed, so I think it is worth saying now.
Everyone uses AI today. But it is already obvious that the same model produces very different results in different people’s hands. This may sound a little rude, but some people still don’t really know how to use agents. They use an agent as if it were just a chatbot. They haven’t set up MCP integrations. They don’t manage memory. They don’t have customized skills. Yet they still describe themselves as “AI native.” Since the idea of AI agents emerged, we have gone through several waves of vocabulary: first prompt engineering, then context engineering, and now harness engineering. The terminology changes, but the underlying idea is remarkably consistent: how do we give an agent the most accurate, comprehensive, and focused information and instructions possible? And that matters because, for the same question, your agent and my agent are very unlikely to have the same context. We use them differently. They have access to different information. Their memories are different. Their tools are different. The instructions around them are different. So the answers we get from them will almost certainly be different too.
That is why the statement, “If I wanted an AI answer, I would just ask AI myself,” is, in many cases, a false premise. Chances are, the person asking you has already asked his/her AI. Maybe their AI did not have enough context to give a useful answer. Maybe it gave the wrong answer. Maybe it could not identify the root cause of the problem. And that is precisely why they came to you. So the fact that your answer ultimately comes from AI does not, by itself, make that answer less valuable.
Your AI and my AI are not the same.
In fact, when what someone really needs is context, a thorough AI-assisted analysis may be more valuable than a short human-written answer. Your agent may be able to synthesize details, history, design decisions, debugging paths, and edge cases that you yourself would struggle to reconstruct cleanly from memory.
I learned this firsthand from a large feature I once built for my team in one of my experience. The project took me about a month. Along the way, I ran into all kinds of problems. I went through repeated testing, redesigns, refactors, dead ends, and what looked like detours at the time but later turned out to contain useful design insights. By the time the feature finally merged, I had accumulated a huge amount of context: debugging history, design documents, abandoned approaches, architectural reasoning, and lessons from things that had failed. I was also the point of contact for the feature, so I knew that sooner or later people would come to me with questions. Rather than wait for that to happen, I took everything I had and asked Claude to turn it into an HTML site: an “all you need to know” guide to that feature. I shared it with the team, and I explicitly said that it had been written with AI. I am fairly confident that another engineer, even with access to the strongest available AI model, could not have reproduced the essence of that document simply by pointing an agent at the merged code. The code did not contain the full story. It did not contain all the failed designs, the debugging process, the tradeoffs, the reasons certain choices had been made, or the context that existed only in my notes and in my head. The document did.
Unfortunately, almost immediately after I shared it, one coworker disliked it. The reason was simple: “This document is entirely AI-written. I’m not going to read a single word of it. I could only laugh. My private thought was: one day, your agent may not feel the same way. Sadly, we will never know. That coworker left the company not long afterward. The two events were, of course, completely unrelated, but the coincidence still amuses me.
I doubt I am the only person who sees things this way. So why did Don’t Paste the AI resonate with so many people? I think the answer is that the frustration behind it is real. The problem is just slightly misdiagnosed. Sometimes I ask a coworker a simple question because I need a quick, human answer. Instead, they send me three pages of AI-generated analysis. That is annoying. But the lesson should not be “never send AI output.” The lesson is that in an AI-native workplace, we need a new communication protocol. The simplest version of that protocol is:
human to human, AI to AI.
Before answering a question, we need to understand which kind of communication is actually being requested. Does the person asking need my judgment? Or do they need information that they are going to pass directly into their own agent? Those are two very different requests. And in the future, I think the person asking the question has an obligation to make that distinction clear. If you need a human judgment call, say so. Then the person answering should investigate the issue, form an opinion, and give you something concise enough for a human to understand, evaluate, and act on. But if what you really want is context that you are going to copy directly into your own agent, then do not demand that your coworker spend half an hour polishing that information into a beautifully summarized human response. Everyone is busy. If someone researches the problem, organizes the context, distills it into a neat explanation for you, and your first move is to paste the whole thing into an agent anyway, then a considerable amount of effort has been wasted.
This new way of working creates another requirement as well. You still have to know your stuff. A person must not only be responsible for work produced with AI. They also need to understand, deeply, what they are submitting. That sounds obvious. But from what I have seen, since many companies started introducing AI agents into everyday engineering work, some people have become remarkably detached from their own output.
You notice it in meetings when someone seems strangely “offline.” You ask a question, and it takes them a long time to respond because they have gone off to query their agent. Or you ask something basic about work they themselves completed, and they cannot give you a clear answer.
Suppose you built a feature and, at a high level, there were two obvious implementation strategies: approach A and approach B. Which one did you choose? That should not be a difficult question. Yet increasingly, it sometimes is.
A lot of people are asking what they need to do in the age of AI agents to avoid being replaced by AI. Being able to understand and explain your own work seems like a fairly low bar.
My goal here is not to attack a particular school of thought, much less any individual person. I am more interested in describing a pattern that is already emerging and making explicit some rules that I suspect many of us already intuitively understand, even if we have not yet said them out loud. So if there is one thing I would encourage, it is this:
When you genuinely have a reason to send someone an AI-generated answer, do not be embarrassed by it. Do not let Don’t Paste the AI become a rule that prevents you from using the medium that best fits the situation.
Sometimes a human answer is better. Sometimes an AI-generated answer, backed by context the other person does not have, is much better. The important question is not whether AI wrote the words. The important question is whether the answer is useful, whether the person sending it understands it, and whether it is the right form of communication for the person receiving it. In the end, practice is still the best test of what works.
I feel that we are undergoing a cultural invasion by AI, an invasion of the whole of human culture by AI. Almost all of AI's training corpus today comes from the accumulated culture, the writing and the history of every human being who has ever existed. But anyone who has actually worked on training large models knows that all of the text humanity has ever produced, the low-value and the high-value alike, stopped being enough for training the next generation of models a long time ago. And yet AI keeps producing text (books, even), images, video and audio without pause. Before long, the AI-generated share of the text on the internet will be as large as the human-generated share. Not long after that, in the internet's entire textual corpus, the writing produced by humans will amount, in sheer quantity, to the very tip of one hair on one of nine oxen. (I do wonder whether there is some Moore's law at work here as well.) AI looks like it is offering choices, but what it is really doing is forming habits. For most of humanity, in the near future, faced with any occasion that calls for formal writing, people will move from not wanting to write, which is where we are now, to not being able to write, and finally to not daring to write. Not being able to write, because AI is quietly consuming the human capacity to organise (complex) language. Once you only have to supply a few keywords and AI will place a mass of seemingly finely crafted paragraphs in front of you, you will stop thinking about word choice and phrasing, about cause and effect, about how a piece opens, develops, turns and closes. Not daring to write, because you are afraid of embarrassing yourself, or afraid that you cannot compete with everyone else who is using AI. If someone else can finish 100 pages in half a day with AI, and it takes you a whole day to write 2 pages, then external pressure alone may leave you with no choice but to use AI, and no courage not to.
I still remember that when I had just started my PhD in 2022, writing a paper took an enormous amount of preparation, most of which was the process of sorting out my thinking bit by bit. Back then, writing a paper was like making a film. Before any real writing began, we would plan out the paper's sections (roughly how many sections, and what each one was for), then plan the paragraphs inside each section (how many paragraphs each section would have), and then, something people may not expect, as a non-native speaker of English I would even design how many sentences went into each paragraph and what each sentence was there to do. It is not as extreme as it sounds. In a standard academic paper, a paragraph of moderate length is maybe five or six sentences. What is being planned here is "what function each sentence serves", rather like deploying troops on a battlefield, and there is even room for a lot of small tricks of the writing craft. For instance, if I needed to praise something before criticising it, I might spend the first sentence lavishing praise on whatever I was about to disparage (usually the baseline, LoL), then lay an ambush (usually in the second sentence), then make the case for that ambush, and so lead into the problem we wanted to discuss. There might be antithesis running across sections, a few puns quoted from elsewhere, a few sentences carrying my own obsessions and my own venting (though smuggling in a private agenda is not a good thing), and before the writing even started the whole thing might go through several rounds of revision. It is just like writing a script and drawing a storyboard for a film. But saying all of this does not mean that what I wrote then was necessarily better than what gets written now that we have AI. Honestly, in large part it probably was not. The difference is that what I was left with was the process of creation, the evolution of the logic, and my own distinctive way of putting things after having thought them through. Who exactly is "I"? I am slowly coming to feel that expression in writing is one of the ways that "I" becomes visible. If every time I speak it has to pass through a proxy, is "I" still "I"? If in the end I cannot even form a complete view and argue it out, is that view still "what I mean"? People in every country in the world hold a deep hostility toward the censorship of writing by central authority, but in the future, will we willingly let AI review, and even rewrite, every utterance we make?
I consider my own writing fairly poor, mainly because I have read so little. To my shame, the number of books outside the curriculum that I genuinely read cover to cover before university was probably under five. The only ones I remember are Lin Yutang's Moment in Peking, Qian Zhongshu's Fortress Besieged and Lu Yao's Ordinary World, all of them read during high school (in middle school and primary school I read nothing else at all). I find it hard to believe now that before university I had not even properly read many popular novels. The result was that my writing was very bad, and writing an essay in every high school Chinese exam felt like being constipated. But now I force myself to keep up the habit of writing, because what makes writing different from speaking is that it shows, more truthfully, the mark a person has left by having lived. When ChatGPT had only just been released in 2023, using AI was something of a fashion, and I was constantly having AI rewrite my content. For someone like me, whose English is poor and whose Chinese is not that OK either, AI was nothing short of a saviour. But going back now to read the blog posts that AI polished at the time, I still break into an awkward cold sweat every so often. These days, for anyone who has never used AI, the output of AI carries the same kind of filter: AI is excellent, AI is complete, AI is advanced, AI is even less prone to mistakes than human beings are. I believed that deeply myself, so much so that even when I had a piece of text that was already close to finished, I would still unwaveringly ask "AI" to "polish" it for me. AI would then live up to expectations and rewrite my article beyond recognition, and only then would I click publish, satisfied. Back then I even felt a certain comfort in this, because my denial of myself had produced a warped expectation instead: the more this text was changed by AI, the better its quality had to be. Thinking about it carefully, the me of that time had already given up the right to express my own views in full. So if this is also happening to most people, how humanity is to recover its cultural confidence is the question I am now thinking about. Can we believe that, compared with AI, language created by human beings is something humans can in fact use to express and to communicate in a better way, a way that shows individual character more clearly and that does more good for humanity's future development? When everyone is using AI, and everyone believes that AI's output is superior, will we still be brave enough to say no to AI's cultural hegemony?
Translated from Chinese by AI
Read in full →More and more mathematical conjectures are being proved or disproved by AI, and my own view of this is a little pessimistic. Whatever the field, and however proud we are of what AI achieves, I do not think that should lead us to dismiss the significance of progress made without it, that is, AI-free progress. Put another way: if a field reaches the point where nobody does the work “the old way, by hand” any more, then over the long run that is definitely bad for the field. The particular problem AI creates for mathematics is this. In future, if someone does make a major mathematical advance by hand, they will no longer be able to prove “I did this without AI”; and once you cannot prove it, the best move available is to fall in with everyone else, since you were never going to get the credit for being “AI-free” anyway. So the “by hand” school in mathematics disappears. What effect that has on the field I am not in a position to judge, but I think it is very obviously a negative one. It is not quite the right thing to say, but there is a faint sense of bad money driving out good. I should also add that what the “by hand” school means, and what its disappearance costs, differ from one field to another. So you cannot rebut the above with “AI accelerated software, nobody in IT writes code purely by hand any more, and IT is thriving all the same.” And besides, is the software industry really thriving the way everyone assumes? We may not know that either.
Translated from Chinese by AI
Read in full →I have had enough of AI-written text, AI-drawn pictures and AI-made videos. I see them on practically every platform: Xiaohongshu, Zhihu, X, WeChat official accounts, even Moments??? No, hold on. I genuinely do not understand you accounts that share papers. What is the point of pasting ChatGPT's raw output, markdown syntax and all, straight onto Xiaohongshu? Are you saving me one call against my ChatGPT quota? Have you even read the thing you posted? And why does a Moments post need AI polishing too? Are you afraid your friends won't understand you? Or afraid you won't win anything at the Annual Best Moments Post Awards?
Translated from Chinese by AI