On Artificial Intelligence
Why building machines that are better than us might make us better.
11 June 2026
Do Not Be Afraid
Welcome back to Consortium.
Technological innovation is the cornerstone of humanity. From the stone tools we crafted roughly three million years ago to the internet that has revolutionised our lives since the 1980s, human beings have consistently made breakthroughs given the resources at their disposal. Every step of the way, each advancement created value by either saving time, saving cost, or saving effort. In a way, every so often, humans have made their lives easier by doing the hard work and coming up with solutions that eventually makes the hard work redundant.
Take for example one of the first inventions that we almost take for granted today — the wheel. Prior to the invention of the wheel, humans employed several methods to move objects — either by rolling them, through sleds and drags, carrying them on their backs, having animals drag them, or through tracks and trails. By inventing this circular shaped stone with a donut hole, all this hard work was cut in half, making the movement of objects and humans considerably more efficient.
Like the wheel, there are several other examples that one can cite from history. The lightbulb saved us the effort of lighting a candle, the telephone saved us the effort to converse with people via letters, and Netflix saved us the effort of going all the way to the theatres to catch our favourite movies. However, in the pursuit to make things more efficient, our innovative capabilities might lead us to a breakthrough that surpasses us, the creators themselves. Enter, artificial intelligence or AI.
Most if not all people today fall into the following personas when it comes to their attitudes towards artificial intelligence —
-
Unaware or fearful of artificial intelligence. This person follows and even purports the argument that artificial intelligence will take their job, will harm them through either misuse or manipulation, or — in some cases — will overthrow humanity in some form of a bizarre uprising.
-
Aware of artificial intelligence products. These people are decently aware of productised AI such as ChatGPT and are mild, moderate or heavy users of these applications to meet objectives ranging from drafting emails to correcting code to personalised guidance.
-
Highly active in the artificial intelligence space. These folks use multiple AI products for a variety of use cases. Some possibly even create new applications with the current technology available as either startup founders or 10x engineers for large conglomerates.
-
Researchers within the artificial intelligence community. These are either academics or industrialists working at the frontier, creating path-breaking technology that revolutionises the way we interact with the world.
A large majority of people fall either into categories 2 or 3, however, there is still a substantial percentage of educated people that believe AI is out to steal from them, hurt them, or remove and replace them entirely. In this newsletter, I unpack all my learnings surrounding artificial intelligence ranging from its applications and technical architecture to where I think the future is headed for businesses and society as a whole.
Happy reading!
Applications
Fundamentally, the fear of artificial intelligence comes from how efficient it is at performing tasks that humans have spent years learning and perfecting. One of the first “impressive” feats that AI was able to achieve was coding. When ChatGPT — one of the foremost interfaces for consumers to interact with AI — burst onto the scene in 2022, its ability to write readable code in various languages and correct errors that humans made was astounding. However, a fundamental limitation that GPT suffered from was that it wrote code and APIs1 that were plausible but incorrect, especially for languages that were newer and packages that required staying up-to-date with latest versions. Moreover, the product failed terribly when used on longer, multi-file projects and thus required a human being to intervene and orchestrate consistently. Regardless of these issues, OpenAI did a nifty job at creating a commercially available product that acted as a pair-programmer for many technical folks to learn new languages, scaffold, unit test, and even speed-run repetitive tasks.
Fast-forward to today, AI has significantly advanced in writing sophisticated code with over 40% of “pushed” code being AI-assisted. In a 2024 report, consulting juggernaut Bain & Company found that generative AI (GenAI) saves about 10-15% of total software engineering time and that improvements of 30% or greater are possible as firms adopt a broader agenda of integrating GenAI at its full potential.
Additionally, this advancement in quality has led to interesting patterns in adoption behaviour from software developers. As a part of Anthropic’s Economic Index 2025 the company studied 5,00,000 coding related interactions across it’s flagship AI product Claude and found:
-
Anthropic’s coding agent (Claude Code) is used for more automation as close to 80% of conversations on the platform were identified as “automation” where AI would directly perform tasks relative to “augmentation” where AI would collaborate with and enhance human capabilities. Relative to this, only about 50% of all conversations with the commercial conversational interface (Claude.ai) were classified as “automation”. As AI agents become more commonplace, more automation of tasks can be expected.
-
The most common programming languages across the dataset are Javascript and HTML and from these, user interface and user experience tasks were the top coding uses. This points to the fact that jobs centered around the construction of user interfaces will face increasing disruption as AI systems get better.
-
Startups are the fastest adopters of Claude Code relative to larger enterprises. Anthropic estimated that ~35% of all conversations on Claude Code related to startup work relative to the ~15% of conversations that were identified as “enterprise related”. Naturally, this is a function of both budget and agility as firms with lesser team members and thus bureaucracy around new software adoption will be quicker at adopting AI tools.
What is most interesting however is that now in 2026, in the race towards building the most valuable AI company, your everyday coder isn’t the only one using AI to augment or automate their work. Leaders of some of the largest technology firms are returning back to their desks and jumping into the trenches. For example, Mark Zuckerburg who started coding again after nearly two decades, has been making edits to Meta’s internal codebase using Claude Code!
Aside from coding, AI has become a reliable assistant in areas such as multifold research, complex writing, constrained financial modelling, and even therapy and companionship. From the National Bureau of Economic Research, Chatterji et al. (2025) published a working paper titled “ How People Use ChatGPT ”. Here are some of the salient results:
Chatterji et al., 2025 — Topics from All Conversations
The working paper’s data is interesting as the most prominent use cases of ChatGPT as of 2025/2026 are “practical guidance” and “seeking information”. As large language models2 (LLMs) now have the capability to conduct research by scouring the internet (an ability they did not possess until three years ago), these use cases make sense. Additionally, all LLMs in 2026 provide polished, human-like answers, and thus humans are more likely to seek advice, guidance, and even emotional assistance from AI chatbots. In a 2025 report, Anthropic covered all the ways in which people use AI for support, advice, and companionship.
Naturally, interpersonal advice (that is typically unrelated to psychotherapy) takes the lead as common concerns and anxieties are what people typically require a “venting partner” for. However, a more interesting finding from this report is that as people chat with Claude about their worries, people tend to express an increasing sense of positivity. This finding suggests that Claude is excellent at not reinforcing or allowing negative thought patterns to take shape — a crucial guardrail especially for the 0.3% of all users who come to Claude seeking any form of psychotherapy or counselling.
Aside from practical guidance and new information, writing forms the next big use case as more people take to posting their views on various social media platforms. However, work-place use cases are possibly what is driving this entire segment.
In the workspace, individuals tend to use Claude primarily for writing — possibly emails, messages, document drafts, etc. Although practical guidance and seeking information are still strong use cases, obviously, technical help is the next large use case for employees — coding, troubleshooting software, financial analysis are all practical examples of this.
Clearly, AI is great at a lot of things. However, no piece of software is perfect. Generations before us have shown us how poorly software (and hardware) can perform, eroding millions if not billions in value. Notable examples include Toyota 2009, Knight Capital 2012, and CrowdStrike 2024. So, what is AI bad at?
The United Kingdom AI Security Institute published a paper in September 2025 titled “ Understanding AI Trajectories: Mapping the Limitations of Current AI Systems ”. Broadly, the paper covers four categorical gaps that AI currently possesses:
-
Performance Limitations : Two leading indicators of poor performance exist. First, hard to verify tasks — AI struggles where outputs cannot be automatically graded (eg. essays or open-ended research). Second, long horizon tasks — here, AI degrades in output quality faster than humans as the length of a task increases. Errors tend to compound as the number of steps continue to increase and LLMs diverge from their original goals.
-
Reliability Limitations : LLMs today make probabilistic errors and also present false answers as fact. Due to these high error rates it is difficult to meaningfully scale AI-led products and services where contracts are high-stakes such as defence, medicine, and law. Additionally, models do not match their confidence to their accuracy thus posing a strong risk if over reliant.
-
Adaptability Limitations : Two fundamental issues here have been that first, beyond the training distribution, no model has been able to adapt to novel real-world settings. Second, when fine-tuned on new tasks, performance on older tasks tend to degrade. Although more recent models have shown less decay, this is an open problem at scale.
-
Originality Limitations : Possibly the more critical gap in current AI capability. Despite the frontier models winning gold medals at the International Mathematics Olympiad, there has been no novel mathematical models or theorems that have come out of AI that are of scientific value. Even AI-written research papers have chased either uninteresting, random hypotheses or have been small variations of existing work.
Therefore, despite scale usage across foundational tasks such as writing, gathering information, designing, and personal advice, AI is meaningfully lacking in being a reliable partner when it comes to long-range, multi-step, complex tasks that involve real-world analysis and original thinking. Within these limitations are where the next set of great companies will be built especially those that create the infrastructure to pave the way for breakthrough capabilities across sectors.
Until now, we have a good grasp on what AI can do well and where it suffers in performance, thus signalling where the next venture opportunities lie. But, how did we get here? And, more importantly, now that we have, how does AI do what it does — whether very well or very poorly?
Mechanics
Today, certain words and terms are heard more than they ever were before. “Training”, “tokens”, “fine-tuning” to name a few. To anybody outside of the tech or private capital industry, these terms would naturally feel foreign. In December 2025 the Searchlight Institute alongside Tavern Research conducted a survey to understand AI literacy in America. With a sample of 2,301 Americans aged 18 and above, the Institute’s online survey contained a wide range of questions on the applications, understanding, and perceptions of AI in America. Three questions in particular stood out as they point directly to how literate Americans and perhaps the rest of the world3 are when understanding the underlying mechanics of AI models.
Across the board, a majority of people either understand AI a “little bit” or “somewhat” with a good ~15% reporting that they do not understand AI at all. The next question then ends up being — what do people understand about AI?
When asked what AI does as people ask it questions, close to 30% of respondents provided an accurate answer, however, almost 50% of people assumed that there is a large database that AI looked up before providing an answer. Additionally, more than 20% of people believed that AI was following a script of prewritten responses. Given this data on AI literacy, the question on most people’s minds must then be — what is really going on at a system level for AI to carry out tasks the way it does?
Teaching Machines to Read
Before a model can realistically provide any "intelligent” output, it has to be able to read. Every text-based character that is entered into an AI system must first be converted into numbers as mathematics is the only language a machine perfectly understands. Thus, the first step in this long process is called tokenisation.
When you prompt Claude, Gemini, or ChatGPT, the system takes your message and breaks it down into smaller components or chunks called tokens. Each token is roughly 75% of a word. For example, the word “running” might become two tokens (“run” and “ning”). Once this is complete, each token is assigned a number from a fixed set of anywhere between 50,000 to 100,000 entries. For example, “Hello” might be token 15497 or “Cat” might be token “4326”. These are simply labels or IDs so to speak that are assigned to each token. There are no mathematical relationships between the token IDs. Moving forward, the model now never sees any letters or words. It simply sees a sequence of numbers.
The next step is embedding. Each token ID is looked up in something known as an embedding matrix. To visualise this better, think of the entire model’s token vocabulary as a giant grid. Each token now occupies one row in this grid and every column in that row is a number. A typical matrix has between 768 and 12288 columns per token.
Imagine each column is a question being asked about every word. For example whether “Dog” is an animal, or if it is royal, or if it male or female. Every word then gets a score between -1 and 1 for each question and those scores fill each cell in the token’s row. The closer to 1, the closer the relationship between the token and the question being asked. “Dog” would score higher on animal but lower on royal and “Queen” would score lower on animal but higher on royal. When taken together across columns, each of these scores represent the position each word takes relative to another word. This set of numbers across columns that are assigned to a token is called a vector.
The critical thing to understand however is that each column header is not written by humans from a list that was pre-decided. Each column header such as “is animal” or “is royal” start off as random numbers. Which raises an important question — how do they become meaningful?
Meaning is derived from a process known as training. Before any model sees the light of day, it is trained on large set of text. This text can range from large portions of the internet, tens of millions of books, academic papers, Wikipedia, and more. As the model runs through this text, it is given a simple task — with the words seen so far in any sentence, predict what the next word will be. The first few guesses the model makes are more often than not incorrect. The system checks for the right answer and then measures how wrong the prediction it made was. This measure is known as loss. The model then takes this loss and flows it backwards through every number in the model, giving each one a small push in the direction that would have made the initial prediction more accurate.
This entire process ranging from making predictions and then pushing the numbers backwards repeats trillions of times across the entire corpus of text that the model is being trained on. Over billions of corrections, the numbers stop being random and eventually start becoming meaningful. It is this very process that gives the embedding matrix its structure. For example, over the course of the training process, the rows for “cat” and “kitten” come closer together as they tend to appear in sentences together while terms like “animal” and “quarterly revenue” never appear near the same words in normal sentences and thus their rows are further apart.
Simultaneously, the column headers also play a role in deriving meaning. As the model is working through its training data, it would attempt for example, to predict what comes after “cat” as it is a living being that can carry out actions such as eating, walking, and sleeping. If through an early correction, a certain column (say, #50) drifted higher for “cat” than for “mat”, this change allows for a lower loss and thus a better prediction for the model. With a smaller loss, the next correction will allow Column 50 to move even higher for living things. Through billions of corrections, eventually Column 50 will end up being consistently high for living things relative to non-living things. Nobody has pre-programmed this. It is entirely a function of the training process. The label “is animal” in the above illustration is a human’s interpretation of what the column is capturing by looking at which tokens are scoring high in this column and which are scoring low.
Something to note here is that the physical structure of the embedding matrix remains the same however the numbers or (“scores” as described above) within the cells tend to move upper or lower (between -1 and 1) and through the training process, similar tokens end up having similar vectors. The final result is that by the end of the training, each token’s row of numbers or vector becomes a precise set of coordinates within the matrix and thus captures its meaning relative to everything else in the language.
The third and final step is positional encoding. As the model processes all tokens concurrently rather than sequentially, it does not fundamentally understand word order. Without a perceivable signal, “cow eats grass” and “grass eats cow” are inherently the same as they are mathematically identical with the same tokens and same embeddings. Thus, a unique numerical tag is attached to each token’s vector based on where it sits in a sequence. A fresh tag is assigned every time a sentence is processed.
Think of it like seat numbers on an airplane.
Therefore, what the model actually receives before producing any sentence is a combination of the embedding vector which carries the token’s meaning and the positional tag which carries its location in any given sentence. This is why the positional encoding takes place after the embedding as the meaning of the vector is required before its position in any sentence can be derived.
Now, the input has been fully translated into a structured representation as each token carries its meaning and position and the model can now reason over this architecture of tokens. The pre-attention phase is complete.
Attention, and Everything After
Now that every token has a meaning encoded as a vector and a tagged position within each sentence, the model faces a deeper problem. On their own, tokens are ambiguous. In the sentence “I sat by the river bank”, the word “bank” means something entirely different relative to when it is placed in the sentence “I deposited money at the bank”. The embedding matrix provides “bank” with a vector but the meaning of the word is largely dependent on the surrounding words. Meaning is a property of words in context with one another.
Thus, to solve for this ambiguity, the attention mechanism was designed. In 2017, Google researchers published a landmark paper titled “Attention Is All You Need” wherein a process is described for each token in a sentence to directly examine other tokens and decide how relevant each token is to its own inherent meaning. Thus, rather than the model reading a sentence from left to right and “forgetting” the first few words by the time the last word is reached, every token is given the ability to look at the entire sentence all at once.
Consider the sentence: “The cat sat on the mat because it was tired.” Here, the word “it” is ambiguous on its own as it could either refer to the word “mat” or “cat”. To resolve this confusion, “it” needs to look outwards and figure out which word it belongs to. The attention mechanism gives the word a structured process to do exactly this through three roles that each token plays simultaneously namely Query, Key, and Value.
Before diving into what these roles are however, it is important to understand that the embedding vector for each token does not get directly used in the attention mechanism. Instead, the embedding vector undergoes a transformation. Aside from the embedding matrix, the model contains three separate matrices of numbers called the query weight matrix , the key weight matrix , and the value weight matrix. These are independent of each other and are tools that are within the attention mechanism. When a token enters the attention mechanism, the embedding vector is multiplied by each of these three matrices, which then result in three new vectors that are ascribed to each role. This can be likened to passing the same piece of information through different lenses, reshaping it for different purposes. Like everything else in the model, they begin as random numbers and get nudged into useful shapes across billions of prediction and correction cycles during training.
Now that process is established, this is what each role does.
The Query is the question each token asks of every other token in a sentence. In the above example, when “it” enters the attention mechanism of the model, its embedding vector is multiplied by the query weight matrix to produce a query vector. For “it”, a pronoun, the query is essentially asking — to which thing do I refer to? The question is not human readable (as it is a vector of numbers), but throughout the training, the query weight matrix is shaped in a way that pronouns naturally produce queries that point towards subjects and living things.
The Key is the advertisement every token broadcasts about itself to other tokens in the sentence. For example, when “cat” enters the attention mechanism, its embedding vector is multiplied by the key weight matrix to produce a key vector. This new vector tells the other tokens in the sentence that “cat” is a living creature, is the subject of the sentence, and is capable of being tired. “Mat” on the other hand has a very different key vector, possibly one that advertises the token as an inanimate object. At this point, the query and the key vectors can be compared to one another mathematically through a measured score. A higher score means a higher relevance between the two tokens. In this example, “it” and “cat” will produce a high score relative to “it” and “mat”.
The Value is the actual information that any token will hand over once another token has decided to pay attention. When “cat” becomes highly relevant to “it”, “cat” will provide a value vector. The information that is provided within this value is meaningful in that “it” requires this from “cat” to resolve its own meaning.
With these roles defined, the entire attention mechanism works as follows. “It” computes similarity scores between the query and the key of every other token in the sentence all concurrently. All these scores are converted into percentages that sum to 100. For example, the similarity score between “it” and “cat” might be 60%, “sat” 20%, and “mat” 5%, while the remaining tokens share the rest. Now that these percentages are assigned, “it” will collect the value from each of these tokens weighed by the similarity score percentages. For example, 60% of what “cat” is offering, 20% of what “sat” is offering, 5% of what “mat” is offering and the remaining of whatever the other tokens are offering in value.
Importantly, every token’s original embedding vector is always preserved and carried forward through addition atop the new contextual information, resulting in a vector that is now a combination of the original embedding vector with positional encoding, alongside weighed information from each of the other tokens in the sentence. With this, the ambiguity of meaning is resolved! “It” is now aware that it belongs to “cat”. This process runs simultaneously for every token in the sentence. Thus, every token asks a question through its query, advertises itself through its key, and collects information from other tokens through their values, weighted by similarity scores. The result of this process is an enriched sentence with defined contextual relationships between tokens.
The attention mechanism runs multiple times. In the original 2017 paper, it runs 8 times simultaneously where each parallel run is using a completely separate array of query, key, and value weight matrices. Each run of the attention mechanism is called a head. Each head specialises in a certain kind of relationship. For example, one head might specialise in resolving the references between the pronoun and the subject, while another might track grammatical consistency where verbs are connected to their subjects. These specialisations are not pre-programmed. They emerge naturally because having multiple heads for multiple relationships produces a better prediction than one head looking for everything. Think of it like multiple employees working towards achieving multiple tasks rather than the CEO running every department. Each head has three weight matrices operating concurrently.
Once the attention mechanism is complete, every token goes through a feed-forward network. The feed-forward step is simply each token digesting the information it has learnt from the attention process. Much of the model’s factual knowledge gets encoded at this step. As the attention mechanism routes and exchanges information between tokens, the feed-forward network allows each token to process and retain this information. The two processes together — the multi-head attention plus the feed-forward network constitute one layer.
Currently, modern AI models stack multiple layers on top of one another with each layer receiving output from the previous one as its input. As the layers increase, so does the depth of analysis from the ground up. For example, the first layer might resolve basic grammatical relationships while the second layer might establish relationships between the words. By the final layers, the model is reasoning about causality, intent, and abstract logic. Each layer builds on what the previous layer has established and progressively becomes richer in the understanding of what the sentence means and which token must come next.
After the final layer is complete, each token produces a probability distribution across the entire vocabulary. Essentially, a score for every possible next token. The model selects the most likely one, adds it to the input and then runs the entire process all over again from the attention mechanism onwards to generate the next token. One token at a time, the entire response is complete.
Understanding how these models work from a systems level perspective is more than just an academic study. Model architecture is fundamental in that it determines where the costs lie, where the bottlenecks are, and where value is created within the larger AI economy. Training a model requires the running of a billion or more correction cycles across trillions of tokens — an exercise that requires large amounts of compute power. Distributing this model to millions if not billions of users requires running the attention mechanism and the feed-forward network for each and every token of every single response. As such, the hardware that trains, the infrastructure that hosts, the tools that allow access, and the applications that are being built, all form one of the most pivotal supply chains in the history of technology. Thus, comprehending the mechanics of the model is crucial for understanding where capital is allocated, where moats are being built, and where the next set of breakthrough companies will emerge.
Supply Chain
Historically, supply chains have formed the bedrock of economic progress from long distance trade via the Silk Road, to the Industrial Revolution enabling the creation of railways, steamships, and aircrafts, to modern ERP systems allowing precise coordination and resource allocation between the various moving parts. What is common across these technological shifts is that each of them had a supply chain that was largely invisible but absolutely indispensable. The artificial intelligence industry is no different. Behind every chatbot, recommendation engine, or autonomous system lies a distinct, multi-layered supply chain that allows these products and services to be accessed by billions across the globe.
Today, 7 layers form the structural architecture of the global AI supply chain:
-
Raw Materials & Fabrication: Layer 1 is the physical foundation of everything. Rare earth minerals are mined, and specialised equipment such as ASML’S EUV lithography are built alongside semiconductor fabs like TSMC and Samsung that print circuits at a nanometer scale onto silicon wafers. This layer forms the bedrock without which other layers cannot exist.
-
Chip Design : Layer 2 consists of companies that are in the news almost every other day. NVIDIA, Google, Cerebras all design processors such as GPUs, TPUs, and NPUs4. These processing units are all optimised for the matrix multiplications that AI models run on. These companies provide the designs of the chips to companies in Layer 1.
-
Data Centres & Networking: Layer 3 takes the manufactured chips and then cluster them into large compute pools. In a practical sense this means that large data centre operators like Amazon Web Services, Microsoft Azure, and Google Cloud provide the infrastructure — buildings, power, cooling capabilities, and connection (eg. NVLink and InfiniBand) to allow thousands of GPUs to act as a single unit.
-
Data, Training, and ML Operations : As we know, data is key to the efficiency that any model hopes to demonstrate. Therefore, before any model can be trained, data needs to be cleaned and then labelled at scale. Layer 4 involves the collection and annotation of data, but also includes the construction of the data pipeline infrastructure, tracking experiments, evaluating the model’s performance, and monitoring its behaviour post deployment. Scale AI, Weights & Biases, Databricks, and Arize are all companies that build in this layer.
-
Foundation Models : Combining the compute power from Layer 3 and the data infrastructure from Layer 4, frontier research labs such as Open AI, Anthropic, Google DeepMind, and Meta train large models and expose them to the world via APIs. This layer is what has unlocked an unprecedented number of applications and companies where every developer can now access top class intelligence without having to own the infrastructure themselves.
-
Developer Tools & Middleware: Model APIs alone are insufficient for production applications. This means that simply being able to send text in and get text back is not enough to build a real product. Real digital products require a model to remember context, pull relevant information from company databases and route requests across different model providers. In order to service these needs, frameworks like LangChain, vector databases5 like Pinecone, and LLM gateways like Portkey all form Layer 6 — possibly the newest and most volatile layer that allows developers to unleash the full potential of AI models.
-
Applications : Layer 7 is where value is finally delivered to consumers and businesses. Within these, we see five buckets — Consumer AI, Enterprise SaaS AI, Vertical AI, Embedded AI, and AI Services.
To better understand how these layers are connected in bringing AI models to life, here is another visual:
What is truly remarkable about the AI supply chain is not just its complexity, but how quickly it has reached this stage of maturity in the short span it has been around. A decade ago, the 7 layers described above did not exist in its current form. Most people remember NVIDIA to be the company making gaming graphics cards! The idea that a software developer in Bangalore or Barbados can access frontier intelligence through simply an API call at cents on the dollar would have been inconceivable half a decade ago. Now, the supply chain is real, robust, and evolving as it transforms not just how software is being built, but how value is being created across every sector in the economy. Thus, the question to be asked is not whether AI is going to change the world, it already has! The question is what a new world looks like, who benefits from it, and what does it demand from those living inside it.
Future of the World
In 2017, I had the opportunity to speak with Mounir Shita.
Mounir is the co-founder and CEO of Kimera Systems, a company that began researching one fundamental question way back in 2005 — what is intelligence? In the pursuit of the question, in 2012, Shita along with Nicholas Gilman started Kimera Systems in Portland, USA. Over the course of three rounds, they raised $378,000 to affirm their central claim that rather than pursuing neuroscience-based approaches to AI, Kimera would use the principles of quantum mechanics to model generalised intelligence6. This thesis developed into the General Theory of Intelligence.
Kimera’s core product was an AGI engine called Nigel. Nigel was centered around the idea that rather than learning from some centralised brain or bank of knowledge, it would learn through a neural network7 with nodes that can extend into any type of connected device. The original application rested on the hope that Kimera would make smartphones smarter by helping them to not only understand what you are doing in real time, but also your motivations behind your actions and thus adapt to you immediate needs.
As an intrigued undergraduate student, I signed up for Nigel’s beta program and received access to an Android version of Nigel that recorded bits and parts of my phone’s activity and would use my voice to run certain functions. Although slightly dysfunctional and buggy, the overall experience was eye-opening — it gave me, at the time — the language articulate to Mounir what I thought AGI would look like while using Nigel.
Re-reading that email today, I realise that my vocabulary is dated but the architecture underneath it is not. In 2017, what I believe I was describing with the words “clock into social media applications”, “reply to messages for me”, and “understand the pattern with which I reply to particular messages in an appropriate context” was what the industry today calls agentic artificial intelligence — an AI system that is autonomous, goal driven, and most importantly adaptable and context aware.
Around a decade ago, I believed that AGI should be packaged as a companion that behaves like a digital twin and adapts to your day-to-day life — something several companies today are in the race to achieve. The collection of nuanced behavioural data and pattern recognition of human beings is yet to be achieved at scale, but is not an unachievable goal in the next decade.
My bet, it turns out, was right. In 2017, I described agentic AI to a founder who was already building in that direction but the industry took roughly 7 years to deliver what I was pointing at. I highlight this purely because ambitious predictions about AI have historically erred on the side of caution and researchers who predicted machines would one day beat humans at chess were called naive. All the engineers who predicted that a device in your pocket would replace your camera, map, bank, and TV were called delusional.
Therefore, with everything I know today about how AI operates and where businesses are currently being built, here are four predictions on what AI will be capable of next:
Understand Human Behaviour & Social Dynamics
Today, AI has been trained on a vast knowledge base of the internet alongside, books, blogs, articles, videos, photos, animations, and other forms of media. It is capable of reading through these and make accurate predictions on what people are looking for when they are either looking for information, writing text/code, or seeking any personal/professional help. AI is also capable of running entire business functions from automating payroll to building financial models. Thus, AI is incredibly good at pattern matching, predicting, and correlating information.
However, AI currently isn’t capable of reasoning on why people behave the way they do. It cannot model other agent’s beliefs, intentions, and motivations, and most importantly cannot reason over truly novel human situations. Although these limitations exist, there has been progress, and thus here is my first prediction:
I. Artificial intelligence will in the next decade have the capability to track, understand, and model over the contextual, contradictory, and path-dependent nature of human beings. With new advancements being made in Causal AI, Mental World Models, and Theory of Mind architectures, the likelihood of AI models being able to absorb culture, trauma, incentives, identity, and irrationality, and allow for these variables to meaningfully interact in order to map human behaviour is plausible.
Understand the Five Senses
AI research is well underway to help models learn from the world around us. Out of the five senses, vision has been solved at a functional level. Similarly, both audio as well as language are at a very mature stage. The remaining three sense are each at different stages in development.
Currently, touch is the most advanced out of the three. Flexible tactile sensing systems are rapidly advancing across various mechanisms, materials, and designs. In 2026, there is a range of visuo-tactile world models that in simple words, are AI systems that can not just see an object but feel it and adapt their behaviour accordingly. Although advanced haptic systems are capable of simulating texture and physical cues through electrostatic modulation8, the next challenge is for touch to be integrated into systems outside of controlled robotics environments.
Smell is also seeing a significant amount of progress. Currently, next-generation AI powered electronic noses are being developed that are mimicking human olfactory senses by converting scent molecules into electrical signals which allow for accurate recognition of diverse and similar scents.
Finally, taste remains as the most nascent research field. Although advancements have been made in electronic tongue systems for on-site quality control, gastro AI in a consumer-facing sense is still majorly theoretical. The challenge is two parts — one, that the hardware for such an endeavour requires molecular contact in a sense that smell does not and two, there are presently no large-scale datasets compared to what has enabled the rapid progression of vision and language models.
Despite this, current progress across the five senses is promising and thus here is my second prediction:
II. Artificial intelligence in the next 20 years will be able to decode and understand the five senses of a human being clearly. World models will advance to the point where if they are installed into consumer or business oriented robotics, they will be capable of accurately predicting and experiencing touch, sight, smell, taste, and sound.
Understand Human Emotions
There is a clear distinction between AI modeling and simulating emotional behaviour versus AI actually experiencing emotions as a state of being which teeters into AI consciousness problem which at this current stage is fundamentally unknowable. Current models are able to exhibit functional emotional analogues in their output. However, whether that is neuroscientifically and philosophically valid is something humans have not figured out with absolute certainty.
Thus, my measured, yet ambitious third prediction is as follows:
III. As artificial intelligence develops an understanding of social cues and human dynamics alongside learn world experiences, it will by sheer consequence, be able to perfectly simulate emotions and develop a sense of sympathy, empathy, and care for other human and non-human agents.
Integrate with Human Beings at Birth
Ample progress has been made in the implantable technology space by companies such as Neuralink and BCI. However, there is a significant gap between the current technology that exists on the market and something that is commercially safe, stable, and scalable such that a person could utilise this technology for everyday use. From a technical, regulatory, and ethical standpoint there are several challenges and thus the “at birth” ambition is a very high bar to clear given the complexities that arise with adults alone.
Regardless, here is my final prediction:
IV. Artificial intelligence will be the core software behind neural devices and assistants and as socio-economic progress is made, regulatory and ethical dilemmas will be ironed out such that scale can be achieved. At birth, children with neural related defects such as auditory loss or slow learning will be fitted with AI-enhanced implants to improve quality of life.
In Closing
Artificial intelligence is omnipresent. The technology behind AI is advancing at a pace most businesses and even consumers cannot keep up with. With new models, new capabilities, and new applications, people are swarmed with options on how to better “optimise”, “improve”, or “enhance” their lives and businesses. However, in a couple of decades, the noise will die down, the dust will settle, and true value will be achieved. The right tools will be built and will be integrated at the right places. Businesses that understand the investment trade-offs will live on while others will perish. People will get better at utilising this technology in ways that meaningfully make a difference for themselves and the society around them. Every piece of knowledge and every tool is at your disposal, the question is if and how you will utilise all of this to shape your world.
If you have reached this far, thank you very much for reading this issue from Consortium. I have loved every bit of research, writing, reading, scrapping, and re-writing. I hope you have enjoyed this as much as I have, and have hopefully learnt something new along the way.
Footnotes
-
An API, or Application Programming Interface, is a set of rules that lets one software program talk to another and request data or actions. It works like a messenger or waiter: one app asks for something, the API carries the request to the other system, and brings back the response. ↩
-
A large language model (LLM) is an AI program trained on massive amounts of text—like books, websites, and articles—to understand and generate human language. By learning patterns from billions of words, it can answer questions, write essays, translate languages, summarize content, hold conversations, and even write code, all without being specifically programmed for each task. Think of it as a computer program that's read so much text that it learned to recognize how language works and can produce natural-sounding responses, just like a person would. Examples include ChatGPT, Google Gemini & Claude. ↩
-
Controlling for global education levels and other socioeconomic factors. ↩
-
A GPU is a processor designed to handle parallel computations on large blocks of data, originally for rendering graphics (games, videos, 3D scenes) but now widely used for AI, scientific computing, and machine‑learning training/inference. A TPU is a custom AI‑accelerator chip designed by Google to accelerate machine‑learning workloads, especially those built around tensors (multi‑dimensional arrays used in frameworks like TensorFlow). An NPU (Neural Processing Unit) is a dedicated AI accelerator designed to speed up neural‑network and other AI workloads, often integrated into CPUs, smartphones, or laptops. ↩
-
A vector database is a database designed to store, index, and search vectors —numerical representations of things like text, images, audio, or products. It is mainly used for similarity search, where results are found by meaning or closeness rather than exact text matching. ↩
-
Artificial General Intelligence (AGI) is a hypothetical form of AI that can understand, learn, reason, and perform any intellectual task a human can, across many different domains. It differs from today’s narrow AI, which is built for specific tasks like translation, image recognition, or chat.
In simple terms, AGI would be a general-purpose mind in software: not just good at one job, but adaptable enough to handle unfamiliar problems and transfer knowledge between tasks. True AGI does not exist yet. ↩
-
A neural network is a machine learning model inspired by the way the brain connects neurons. It learns patterns from data by passing inputs through layers of connected units and adjusting the connection strengths during training.
In simple terms, it takes input, processes it through one or more hidden layers, and produces an output such as a prediction, classification, or generated result. Neural networks are widely used in image recognition, language processing, and other AI tasks. ↩
-
Electrostatic modulation means using controlled electric fields on a surface to change how touch feels, usually by altering friction or adhesion as a finger moves across it. In haptics and touch simulation, it lets a system create sensations like roughness, stickiness, or texture without moving mechanical parts.
For AI world models that simulate touch, electrostatic modulation would be the part of the system that converts a predicted tactile interaction into a tunable electrical signal for haptic output. In other words, the model estimates what the surface should feel like, and electrostatic modulation is one way the hardware can render that feeling to the user.
A simple example is a touchscreen or haptic panel that increases electrostatic attraction when you slide a finger over it, making the surface feel more resistant or textured. ↩















