In recent months, the news about artificial intelligence (AI) has been overwhelming, making me feel like a resident in the Three-Body World1 watching the sunrise on the horizon, knowing that something significant is about to happen, but uncertain whether it will be one or two suns rising. At this moment, personal contemplation may not be particularly useful, but it is instinctive for humans to reflect. Therefore, I am writing down some of my basic thoughts here, for the purpose of sparking further discussion. It should be noted that this article is not an introduction to widely accepted scientific research results, but rather represents my personal opinions. There are probably mistakes and limitations in my understanding, and all comments are greatly welcome.
Why Is ChatGPT a Critical Point?
Many people believe that the emergence of ChatGPT is a significant event, possibly leading to a technological revolution in AI. However, what kind of transformation is this exactly, and what criteria should be used to assert its importance? For instance, the invention of agriculture and the industrial revolution were milestones in human civilization in terms of energy utilization. By contrast, the AI revolution is clearly about information rather than energy. This point has been reiterated time and again, but here, I wish to delve deeper into the question of what kind of critical point the birth of large language models (LLM) represents.
I believe that what is important is not how much ChatGPT can achieve or how much more it can do than previous AIs. The crucial issue is complexity - what we see today is the critical point of AI's information processing complexity.
Let us first discuss what information processing is. From the perspective of information, all events in human history are essentially processes of information processing. For example, the invention of the wheel required someone (or many people) to come up with the idea, develop the production method, and then pass on this information. After the invention of the wheel, people were genetically no different from those before, but they could enjoy the convenience that the wheel brought because this information was passed down. Zhang Liang's "planning within encirclement, deciding victory a thousand miles away" is information processing, and his messenger passing on this command to the frontline is also information processing. Since the term "processing" gives people a passive and static impression of information, I believe that a more accurate description of the spread and evolution of information is "information dynamics." Human history is itself a process of information dynamics, and individuals are the medium and carrier of the evolution of information dynamics. Different information processing has different complexity, such as Zhang Liang's decision-making complexity, which is obviously higher than the complexity of the messenger passing on orders. Roughly speaking, complexity measures the minimum number of simple steps required to accomplish something. This complexity has three aspects: the complexity of input information, the computational complexity of information processing, and the complexity of output information. If Zhang Liang faces complex and confusing information from the front line, considers all aspects of the problem, and finally instructs the messenger to send a single word "retreat," then in this example, the complexity of input and processing is high, and the complexity of output is low.

Figure 1. Information dynamics before and after ChatGPT.
Now let us return to the present day. In the modern era, prior to the emergence of Artificial Intelligence Generated Content (AIGC), the inventions of computers, the internet, and smartphones have already greatly transformed human society. However, from the perspective of information dynamics, one can observe that the tasks performed by machines are very simple and can be categorized into three types: copying, pasting, and sorting. The internet and social media have made information propagation (i.e., copying and pasting) between people more convenient. Google search and various recommendation algorithms are sorting tasks. The information dynamics of this era can be represented by the network shown in Figure 1(a). From the perspective of complexity, it can be said that all complex information processing happens in the human brain. Different from copying and pasting, sorting does involve a complex computation, but its output is ultimately simple, just a sequence of numbers. Therefore, after this complex operation (say by Google search or a recommendation algorithm), machines can only return this sequence of numbers to humans without being able to use their computational capabilities to iteratively process and accomplish more complex tasks. When a human reads about an event and expresses her opinion in a twitter post, this task already involves more complex information processing than any machine before ChatGPT.
Of course, prior to GPT, Artificial Intelligence has already demonstrated impressive performance in specific tasks such as Go and painting. However, in terms of information processing, these AI models have the same problem as sorting algorithms: the output form of information is relatively simple and fixed. Therefore, this kind of information processing capability cannot be iteratively used to achieve higher complexity. In other words, the complexity of information dynamics depends on the bottlenecks of input, processing, and output.
The revolution brought about by ChatGPT is precisely because it has achieved human-level complexity in all three of these steps. GPT can accept fuzzy instructions in various languages and programming codes. It can provide reasonable output in response to various inputs, and the output itself is also as complex as human-generated content, such as writing articles and code. After the emergence of GPT, the information dynamics diagram is transformed into Figure 1(b). It is known that even the latest version of GPT-4 still has many problems. It often makes mistakes or even nonsense in unfamiliar situations. However, in terms of complexity, I believe that it has already reached the level of human performance. The key issue is complexity rather than specific performance in a task: once the complexity of information processing reaches human-level, the fundamental obstacle to accuracy and other aspects is removed. Given the speed of AI evolution, once it reaches human-level, it will soon surpass humans.
Why the Language Model?
Because complexity is key, the AI revolution emerged from large language models, rather than other AI fields.
The first sentence of the Bible says "In the beginning was the Word," and the first sentence of the Tao Te Ching says "The Tao that can be told is not the eternal Tao." The emergence of human language was a pivotal event in the history of Earth, marking the shift from genetic inheritance and mutation to the spread and evolution of language as the most important information dynamics on the planet. From an information dynamics perspective, the rapid evolution of AI since the advent of ChatGPT may be the second most important event on this planet after the emergence of human language, marking a shift in the most influential information dynamics process from chemical processes in the human brain to electronic processes on chips.
Why is language so important? Human language is not just a tool for transmitting fixed signals like the dances of bees, but rather a general tool that can be used to describe anything from concrete to abstract. We can not only talk about objects in the real world, but also describe the relationships between them and the relationships between relationships. In the real world, there are only apples, oranges, and bananas, but humans can create the abstract concept of "fruit" from them. Different concepts such as fruit and vegetables also belong to the broader concepts of plants and nouns. These different concepts belong to different levels. An AI system for image recognition can learn to identify the concept of "fruit" from specific images through training, but to learn the next level of abstraction that fruit and vegetables belong to plants, one has to train a separate model. The magic of language is that once we see all of these concepts as words, they all exist on an equal footing. Whether it's "apple" or "plant" or "nonlocality in quantum mechanics," they can all become objects of thought. With language, the world in our mind is no longer just a reflection of the external world, but has a new dimension with infinite possibilities. With this new dimension, the structure of the world becomes flat, and abstract structures that were once layered one on top of the other become as accessible to us as an apple. Using language, we can understand the concept of a straight line or a triangle, summarize the axioms of Euclidean geometry, and use them to prove the Pythagorean theorem. Once all right triangles have been proven to satisfy the Pythagorean theorem, we can master and apply this knowledge without the need for any additional data.
The limits of language are not the limits of human ability, but they are the limits of conscious thought. Humans can learn some skills, such as riding a bicycle, through training that does not require language or thought, but without language, such skills cannot be improved through thought or spread through communication. For example, we can write a bicycle riding manual, but reading the manual will not teach us how to ride a bicycle; we must learn through practice. Therefore, the complexity of the world that we can understand and communicate to others cannot exceed the scope of language. That’s why Wittgenstein wrote, "The limits of my language mean the limits of my world."[1] In other words, for humans, information dynamics are language dynamics. This dynamic includes rigorous deduction and reasoning, as well as leaps of inspiration, even daydreams and hallucinations.
Because of the central role that language plays in our world, it also has a unique position in the development of AI. When we look at the development of ChatGPT, it is not difficult to imagine that a language model could one day learn to drive autonomously, but conversely, an AI designed for autonomous driving would have a hard time learning language. Ted Chiang said that ChatGPT is a "blurry jpeg of the web" [2], which I think is somewhat true, but this metaphor is too static. More important than static knowledge is the dimension of time: ChatGPT is a fuzzy impression of human language dynamics. In other words, ChatGPT has not really learned to think, but it has learned to roughly mimic human thinking processes. For example, mathematician Terence Tao has shown how ChatGPT can suggest proof strategies for unknown theorems, and even though its proof has errors, it can provide useful inspiration[3]. ChatGPT can make such suggestions because, although the theorem is unknown in mathematics, ChatGPT knows how to apply the strategies or patterns of other proofs that it has seen before to this theorem. There is a common view that AI can only imitate and not create, but I believe there is no absolute divide between imitation and creation. In fact, some of the most creative ideas of humans are not completely original, but rather come from analogies based on known thought patterns and knowledge. Newton's idea of gravity from seeing an apple fall is also an analogy of the known and unknown. This analogical process is not essentially different from the attempts ChatGPT made for theorem proving.
As an example, I asked ChatGPT to speculate on some unexpected breakthroughs that might occur in the future study of quantum gravity, and below is the answer it gave me. Although this cannot be considered a particularly exciting idea, the direction of speculation is somewhat reasonable. We can say that ChatGPT is not weak in terms of its ability to brainstorm, and it may even be stronger than humans because of its vast knowledge. However, its problem lies in the fact that it cannot independently verify which direction is more feasible and accurate among many ideas. This is known as the “grounding” problem.

Figure 2. A GPT-4 answer to a prompt asking for an unexpected but reasonable idea in future quantum-gravity research.
Can AI Truly Learn to Think?
The grounding problem, i.e. the difficulty ChatGPT is having in distinguishing facts with hallucination, is deeply related to the fundamental difference between current day LLM and humans. Clearly, GPT-4 is not Artificial General Intelligence (AGI) yet. What is the key thing that is missing?
Daniel Kahneman, in his book "Thinking, Fast and Slow,"[4] points out that human thinking consists of two systems. System 1 is our fast, intuitive, automatic, and unconscious way of thinking. It processes most of the tasks in daily life, such as recognizing objects, expressions, language comprehension, and making simple decisions. System 1 is often based on experience and achieves fast decision-making through association and pattern recognition. However, this kind of fast decision-making is often susceptible to cognitive biases. System 2 is our slow, analytical, conscious way of thinking. This system requires more attention and effort to operate because it handles complex problems, logical reasoning, planning, and long-term decision-making. System 2 can correct System 1's errors, but it operates at a slower speed.
Today's large language models are essentially a simulation of System 1. They directly output text based on an input of words according to a probability distribution, which is similar to humans making judgments based on intuition. For example, with a math problem, GPT-4 can provide a deduction process according to your instructions, but if you ask it to give the answer directly, it does not conduct the deduction process and gives the result directly based on the probabilistic next-token prediction, which I consider as the analog of “intuition” in the System 1 of humans. That's why when GPT-4 is given the instruction "write out the deduction process," its calculation accuracy will significantly improve[5][6]. From this example, we can see that GPT-4 has learned to use language but only to communicate with humans, not to think consciously with language. Using language to think is the main difference between System 2 and System 1.

Figure 3. A schematic of artificial intelligence with System 1, System 2, and memory.
To make language models learn to think, two prerequisites are necessary: (1) they need to have long-term memory. Current GPT-4 has a certain memory of the context of the conversation, but these memories are cleared after the conversation ends. Although it "remembers" a large amount of knowledge, those are not memories acquired during the conversation. If we compare it to humans, GPT's knowledge is more like innate abilities in humans, such as babies knowing how to cry and suckle. To enable language models to learn from experience like humans do, they must have long-term memory of their own history. (2) Language models must be able to process long-term memory and learn from experience. When humans make errors in solving mathematical problems and subsequently learn the correct solutions, they update their long-term memory to incorporate the accurate methodology. As a result, the likelihood of repeating the same mistake in the future is significantly reduced. This learning mechanism enables us to process and retain information with minimal data input, making it far more efficient than the training of AI models. Although GPT also corrects errors, it will still make the same mistake again[6] next time unless the model's parameters are changed through further training.
Recent research endeavors have aimed to equip language models with long-term memory and the ability to access it. For instance, an architecture known as "Reflexion" has been introduced[7], in which the method GPT employs to solve a problem is documented and forwarded to another language model. This secondary model then "reflects" on the recorded process and offers feedback on how GPT can enhance its approach in future attempts. By incorporating this feedback into the prompt, the performance of GPT on certain tasks have been shown to improve by as high as 30%. In another work, the authors designed a virtual town with 25 AI characters, each with their own personality and memories.[8] For example, when two characters meet for the first time, this event is recorded in their memory and they will remember it the next time they meet. When encountering a new event, a character will search for the most relevant memory from their memory as a reference to decide on their current behavior. This architecture enables complex interactions between multiple characters (such as planning and organizing a birthday party). As a third example, the popular program "autoGPT"[9] allows GPT to first plan a complex task and then use various resources on the user's computer to execute it step by step. All of these efforts are aimed at giving artificial intelligence a system 2, like that of humans, with its own "mental activity" and the ability to continuously learn and improve in a context. The structure of artificial intelligence with both system 1 and system 2 can be briefly summarized in Figure 3. I believe there will be rapid progress in the near future along the direction of exploring such “System 2” architectures, even without further increase of the complexity of the LLM itself. This is because, from the complexity perspective, the information processing complexity of this system 2 is not significantly higher than that of the current language model's system 1.
Science Research in the Era of Artificial Intelligence
AI has the potential to bring about significant changes in scientific research, and as a physicist, the author is naturally interested in exploring this topic. From the perspective of information dynamics, scientific research is also a process of input-processing-output, but what sets it apart is its creative aspect: the goal of scientific research is to output new knowledge that did not previously exist. The scientific community of researchers operates like a neural network, where the output of each work becomes the input for future works.
To create new knowledge, researchers must first digest and understand existing knowledge, which can come from sources such as books, research papers, and accumulated practical experience. The ability to learn useful knowledge despite various obstacles is an essential quality of capable researchers. Moreover, understanding the same thing from different angles or explaining it in different languages can lead to a deeper understanding and the creation of new knowledge. Researchers use their understanding of existing knowledge to shape their ideas gradually, similar to how a vague drawing in Midjourney becomes clearer over time. After completing a project, a crucial step is to disseminate the new knowledge, which can be achieved through writing papers, giving academic talks, etc. The way in which information is disseminated can greatly affect the impact of a research project, which is why attempts in getting papers published in top-tier journals become an important and time-consuming task for many researchers.
With AI's human-level information processing ability, each of these stages in the scientific research process could be significantly altered. In the information input stage, AI can assist human researchers in understanding other authors' papers more quickly and thoroughly by providing summaries of various levels of detail based on researchers' needs. AI can also leverage its vast knowledge to point out what knowledge may be useful for current research in unfamiliar fields. AI can make the output of research more flexible. For example, while listening to talks or directly discussing with authors is often more efficient than reading papers, researchers do not always have the opportunity to communicate directly with authors. If AI can explain papers like authors and answer questions, acting as an author's agent, it would be very helpful for expediting scientific research.
In the output stage of information, such flexible output could fundamentally change the way scientific papers are published. Instead of publishing a paper that is a fixed text, the authors could “teach” their ideas and results to AI, and the AI could output these results to the reader in a multi-module form. The readers could choose to learn this work in whatever way she prefers, ranging from a five minute introduction, half hour lecture, to a written paper. Publications can take the form of "alive" knowledge carriers that the readers can interact with.
In the process of creating new knowledge, AI can also propose new ideas and problems and suggest possible directions for attempts based on past experiences. This is something that the current ChatGPT can already try, but as we discussed in the previous section, more accurate understanding is needed to make its suggestions more valuable. I believe that future scientific research should be "AI in the loop," with AI involved throughout the entire process, from transactional work to creative work, making the information processing of the entire research activity more efficient.
More importantly, AI will not only change the way each independent research group works but will also bring new possibilities for collaboration between different groups. In some fields such as experimental particle physics, research has already evolved into large-scale collaboration, with possibly hundreds or thousands of authors on a paper. However, in most areas of fundamental research, cooperation is still limited to a few people or a dozen people. In particular, in the basic theoretical research that I am engaged in, assuming that all scholars have unlimited funding and can expand their own group size as they please, it is unlikely that each person will bring many more students than they do now, and the depth of collaboration and communication between different groups is unlikely to differ substantially from now. This is because, in terms of producing original results, the bottleneck is not in resources (although it is not possible without resources - funding agencies, please do not cut our research funds after reading this article), but in the time and intellectual costs of high-quality information processing. If the research group is too large or has too much cooperation with other groups, the time needed for everyone to understand each other's ideas may take up too much time, resulting in a loss outweighing the gain. Therefore, the discovery of an important idea in reality often rely greatly on small probability events, such as two people from different fields unexpectedly meeting each other and creating sparks. Consider two researchers, each has a valuable but vague idea and shared the idea with people around them. What often occurs is that one of them talks to the ``right” person and they inspire each other in discussions, leading to an outstanding work, while the other talks to someone who is uninterested in this idea, and the discussion ends up going nowhere. The emergence of AI will not change these contingencies, but it will make the entire trial process much more efficient. AI could help researchers surpass the limitations from their environments, so that all researchers can access a more inspiring and stimulating research environment that was previously only available in a few top research institutions.
Perhaps some people may feel that this prospect is too dangerous. If AI surpasses human intelligence, will we researchers lose our job? Possibly, but despite this risk, I still look forward to this development rather than fear it. Imagine if someone gave Isaac Newton a time machine, so that he could travel to the 20th century and learn about modern physics, I would guess that he would probably be excited rather than feel disappointed that he won’t have a chance to discover quantum mechanics and relativity himself.
Potential Crisis
Undoubtedly, the swift emergence of AI carries significant implications for society. The future could manifest as either a harmonious Stable Era or a tumultuous Chaotic Era. Open letters have emerged, advocating for a six-month hiatus on the development of colossal AI models, to enable humanity to better prepare[10]. While such concerns are valid, I believe it is too late for a pause. The development of large-scale models does not rely on clandestine methods, and even if OpenAI ceases operations, others will promptly create similar models.
In response to potential crises, humanity's focus should not be on slowing down AI development, but rather on becoming acquainted with it and learning how to harness its immense potential. We must contemplate how society can adapt to the AI era, establish effective resource allocation mechanisms, and construct a safety buffer between AI and crucial social functions. It is impossible to be fully prepared for every challenge today, but it is essential to initiate societal discussions as early as possible.
I am convinced that AI will ultimately resolve the issues it introduces, much like the social problems brought about by the industrial revolution could not be addressed by eradicating industry, but by utilizing the wealth it generated in a sensible manner. Once AI develops a profound understanding of human psychology and society, it can also assist us in simulating potential future societal issues and experimenting with possible solutions. There are many different aspects in the discussion of AI safety, and I will only discuss one particular aspect in this article.
One of the biggest differences between AI and humans is that the "brain" of AI can be made more transparent than the human brain. This is not to say that we understand the computational process of AI - the interpretability of AI models itself is a difficult research topic - but that when AI performs a multi-step task, the process of "thinking" in language (i.e., the use of System 2 as discussed earlier in this article) can be recorded and shared with humans. In contrast, human mental activities are inherently confidential to others. Of course, the thinking and execution process of AI can be encrypted, making it only accessible to its owner. I believe that an important principle for ensuring the safety of AI is to require that powerful AI models must remain transparent, in the sense that its "thinking process" can be checked, interrogated and regulated by the public, rather than becoming a black box completely controlled by a small group of people. In practice, how to ensure such transparency while maintaining the security of private data is an important technical challenge.
Conclusion
Our generation has long recognized that we live in an unparalleled time; however, it is only now that we have genuinely entered the dawn of this era. The future is laden with considerable uncertainties, but I am inclined to believe that, “Subtle is the AI, but malicious It is not.”
References
[1] Tractatus Logico-Philosophicus, Sec. 5.6, Ludwig Wittgenstein
[2] https://www.newyorker.com/tech/annals-of-technology/chatgpt-is-a-blurry-jpeg-of-the-web
[3] https://mathstodon.xyz/@tao/109971374075988443
[4] Thinking, fast and slow, Daniel Kahneman, Macmillan, 2011.
[5] OpenAI, “GPT-4 Technical Report”, arXiv preprint arXiv:2303.08774 (2023)
[6] Bubeck, Sébastien, et al. "Sparks of artificial general intelligence: Early experiments with gpt-4." arXiv preprint arXiv:2303.12712 (2023).
[7] Shinn, Noah, Beck Labash, and Ashwin Gopinath. "Reflexion: an autonomous agent with dynamic memory and self-reflection." arXiv preprint arXiv:2303.11366 (2023).
[8] Park, Joon Sung, et al. "Generative Agents: Interactive Simulacra of Human Behavior." arXiv preprint arXiv:2304.03442 (2023).
[10] https://futureoflife.org/open-letter/pause-giant-ai-experiments/
Footnotes
-
A hypothetical planet orbiting a chaotic three-star system, described in Liu Cixin’s trilogy The Three Body Problem. ↩