Showing posts with label LLMs. Show all posts
Showing posts with label LLMs. Show all posts

Thursday, June 25, 2026

Note on Large Language Models

LLMs have a certain analogy to compression. From the training data D we obtain a LLM T(D) which is supposed to contain (or "extract") the essential "information" or "statistical patterns" present in D. T(D) is much smaller than D. It is speculated that Claude models are trained on D of the size of a petabyte and that the models themselves range from 150 to 500 GB. The response to a given prompt is analogous to decompression. Supposedly T(D) can "generate" an approximation of all the information originally contained in D. Some questions:

1. Is it not true that the passage D -> T(D) is not lossless, that important information present in D is lost in T(D) and cannot be recovered by it?
2. Is there any way to study T(D) as a mathematical object, detect its structure and geometry? And to study likewise the correspondence between D and T(D)? If there are limitations to doing this are they practical or theoretical?
3. There is an analogy between passing from D to T(D) and passing from general to countable models of ZF set theory (which exist by the downward Löwenheim-Skolem theorems)?
4. Is there not some analogy between forcing using countable models and generic sets and the process of training to generate T(D)? In both cases there is pattern generalization from fragmentary data.
5. Is there any structural correspondence between the structure of T(D) and structures found in the world (not counting neurological analogues of MLPs)?
6. Can we construct toy universes, toy languages and toy training data and study how D -> T(D) works in this simplified idealized scenario to gain more insight regarding real world LLMs?
7. Do LLMs express an essentially emergent phenomenon in which hardware capabilities are a crucial factor? Can we formalize rigorously such a concept of emergent phenomenon or capability?
8. But most importantly LLMs are linear statistical predictors (next token predictors) and they are trained as such. We need to formalize clearly what LLMs are supposed to do in the first place. Suppose we have a (first-order) model M that represents the world. We want our LLM T to be able to deal with a good degree of approximation with the theory of M, Th(M). We are given a finite large set L of first-order formulae with probabilities of their belonging to Th(M). A transformation is applied to L to obtain the object T which is able to include the reliable part of L in Th(M) and to extrapolate to other elements of Th(M). Is this to be understood as both logical and statistical inference?
9. A LLM is just a finite state automaton. But recursively axiomatizable theories are in general not recursive. Can we can construct a theory T such that for any finite subset L of T all LLMs trained on L will err to an arbitrarily with regards to infinitely many sentences of T. We define metrics on expressions, that's the key.
10. And most importantly: are LLMs analogous to syntactic (and algebraic) models used in logic and category theory?  Or the training data is like the a poset P with the dense topology and the LLM is like the topos of sheaves over this site?

11. Do LLMs function essentially by analogy, metaphor, induction and extrapolation? This is an old idea in AI. 

Saturday, May 23, 2026

Yet more short considerations on AI

The philosophical significance of LLMs and a potentially powerful philosophically based critique of LLMs have not yet been developed and perhaps their importance has not even been recognized yet.

Could LLMs be tied to a certain philosophical view regarding the mind, language and the world? And if such a view is manifestly erroneous could not this furnish a sound basis for acknowledging - alongside numerous other reasons - the social and cultural harm of LLMs?

For instance, could we not explore the relationship between LLMs and Quinean extensionalism and meaning-as-use theories? Are not LLMs based on the rejection of the irreducible intensionality of human symbolic activity? Behind every symbol there is an intension. And a formal theory of intensionality must itself acknowledge the intensions of its meta-symbols.  But there is no intensionality in LLMs beyond that of the humans involved in their creation. Searle's Chinese Room ignores the deeper philosophical meaning of computation - or a potential associated geometric theory of meaning - but may be of interest for a critique of LLMs.

Do LLMs compute? Computation is a primordial intensional human activity (see our paper Analyticity, Computability and the A Priori). 

Perhaps there are other machine learning models of language of a more geometric or even combinatorial-algebraic nature which would have interested Riemann who wanted to relate meaning to a kind of cognitive geometry.   Perhaps statistical regularities should be traced back to geometry.

What are LLMs really, formally? Can we formalize the critical values wherein they become 'adequate' for their proposed task? What exactly in 'large'? How can this be formalized rigorously? Can LLM techniques be used for formal axiomatic-deductive systems and automated theorem proving? Or can we prove certain fundamental limiting theorems about the powers of LLMs akin to the unsolvability of the halting problem and Gödel's incompleteness theorems?

LLMs only exist because of the Internet. The Internet and LLMs are part of the same historical-cultural-technical process.  This ontological process might be described as the datafication of humanity. Language ceases to be a tool of human thought, communication and culture-building but rather a tool for the reduction, degrading, emptying, perversion and commodification of humanity itself. The Internet and LLMs are the anti-Gutenberg. Man has become text, sound, image, data, statistics. LLMs are a counterfeit reality, a monstrosity, the world becomes one big corporate controlled screen.

 If we compare the training data and the resulting LLMs is there or not a loss of information? Would it not be more worthwhile to develop sophisticated search algorithms and querying language to access the training data directly?

 LLM culture is the culture superficiality, atomization, banality, cosmetics and deception (a LLM is almost a trained deceiver in the biological sense - it detects and mimics human patterns). There is a loss of the multiple layers of meaning behind every symbol which cannot be reduced to statistical correlations with other symbols. Meaning is replaced with arbitrary social-statistical emergent correlative patterns.

The realm of pure mathematics - and that of pure logic, combinatorics and computability - is a pure realm which LLMs cannot touch or corrupt.  So the formal mathematics and formal philosophy projects, contrary to popular misconception likely resulting from deliberate propaganda and deception - are the antithesis of and antidote for LLM culture. There is a pure universal computational-mathematics-akin language (far beyond the natural language or the audio-visual data  that can be perverted and imitated by LLMs) and a pure logic and a pure thought and mankind may indeed hope to attain them. 

LLMs are not intelligent and statistics cannot solve formal computational problems nor can they encompass the pure a priori synthetic principles of formal computational systems (i.e. the cognitive certainly of the foundational principles for metatheoretic knowledge) as detailed in our paper mentioned above.

No statistical pattern analysis of the shadows are sufficient to lead to knowledge of the object. Pixel injection in image recognizers demonstrates this fact, and similarly for formal reasoning.  Besides lacking intensionality, LLMs lack reference and context (despite the misleading terminology of context windows). 

And most important of all since the true intellect is inseparable from morality, empathy and compassion, completely beyond the reach of LLMs. LLMs do not have the bondage to an illusion of a self.

A major task of philosophy is to destroy the evil empire of LLMs and a good starting point is deconstructing and refuting the worldviews (extensionalism, meaning-as-use) which LLMs embody. 

Is a lawyer someone trained in a specific system of laws or someone who has developed the skill to study, interpret and apply any given system of laws or perhaps be able to cope with significance changes in present laws?  Such a metalawyer is the analogue of a universal Turing Machine. A LLM could never be a judge or a metalawyer - it could not grasp the spirit of legal institutions or the deeper meaning of a given legal context.

Billions use light-bulbs without understanding the underlying physics. Billions could use LLMs thinking they are conscious.

LLMs will become more interesting in the measure in which the multilayer perceptrons models are replaced by geometrically and mathematically more sophisticated models (KANs are a step in the right direction).

Saturday, May 9, 2026

Short philosophical considerations on AI

Hegel and Heidegger were thinkers about their own time, thinkers about historical events and happenings. Few have the insight and courage to fathom the full depth of the meaning of an historical event, the coming to be (coming of age?) of an historical process. Tragically,  it is only some time after the event (the time of monsters?) has hit humanity with full force that Hegel's famous owl can spread her wings. Is it not true that some of what the prophetic author of Sein und Zeit wrote about technology only makes full sense at the present?

This seems to us particularly true of the emergence of the age of the Internet and the age of generative AI which is its logical development.

And yet nothing could be further from our own philosophizing than any form of historicism or historical philosophy. As such both Hegel and Heidegger, for all their interest and insight, must be considered as having crafted systems based on an incorrigible error. 

Social progress is not a law of nature but a legitimate hope - even if at present it seems a distant one - and it is our moral duty to work towards it in the midst of uncertain outcomes.

The advent of the internet was the advent of connection between people. This connection carried rhetorically moral undertones and echoed enlightenment ideals about the desirability of sharing and making knowledge available.  In the present age of the generative AI based Internet powered and controlled by corporations and governments aligned to anti-enlightenment ideals, it may be that it is morally called upon us to practice instead the process of disconnection and the purification and preservation of knowledge(not obviously in the sense of the 'great simplification' of the Canticle of Leibowitz).   

The most basic step is ensuring locality of core information. That is, to be in possession of machines onto which have been downloaded significant portions of Internet Encyclopedias (despite their serious shortcomings) as well as some decently performing LLM. To this we add, it needs not be said, massive of digital preservation of human cultural artifacts, notably libraries.

One can use Kiwix and download for offline viewing the most recent English Wikipedia (50GB text-only 150GB with pictures).  With a AMD Ryzen 7 5825U processor with 16GB RAM and 2GB Radeon Graphics one can use Ollama and download and use some decently performing LLMs (gemma4 comes in E2B, E4B, 31B and 26B A4B). 

Most living beings alternative between states of being awake and of sleep. Can would we design a dynamic LLM which similarly alternates between states of user interaction and re-training based on this interaction? The most important being the correction and/or updating of knowledge or perhaps the removal of harmful and biased content and "thought patterns". If LLMs can improve then it is not only  a question of having the number of parameters equal to the number of neurons of the human brain.

Computers can enhance and aid human cognition as well as hinder and destroy it (there is a growing body of evidence concerning the disastrous effect of excessive or inappropriate generative AI use for individual mental health and cognitive development, not to mention for society as a whole). 

But the harms of generative AI have little to do with lesser-known extremely powerful and beneficial aspects of the computer for human cognition. We cannot go into this in detail here. Let it just be said that it involves using adequate software for the rigorous formalization of human scientific theories and concepts (specially logic and mathematics and formal methods in the sciences) and the vital feedback-loop between human thought and the software interface (IDE) which results in the simultaneous enhancement of human understanding and production and the quality of the software-based formalization and implementation. 

The software in question includes not only Rocq (formerly Coq), Agda and functional programming language but such languages as Python, Javascript and C/C++. Python is a multi-paradigm and highly versatile language with an elegant syntax. Python comes close to achieving the ideal of a universal language in the Leibnizian sense and is a wonderful tool for formalization, implementation,  verification and exploration in a variety of areas in mathematical logic and finite mathematics. 

We note also the importance of minimalism (using as few dependencies as possible) and building things from the ground up - this goes for scientific and philosophical projects, not of course for commercial and industrial ones.  We will address in the future the question of the possible role of machine learning in this process.

Monday, April 27, 2026

On generative AI

Is generative AI corrupting human knowledge and language and by extension human thinking and human culture themselves?

A wikipedia dump is around 100 GB. Wikipedia could be improved and be semantically formatted to be computer readable and advanced query systems could be developed. Would not this be better for the acquisition of knowledge and the advancement of science? Are AI generated summaries of books or papers valid replacements for human ones? What justifies our trust in generative AI as compared to a search engine?

Generative AI is corrupting the internet. Maybe it is a zombie or Frankenstein of human language and knowledge. Or a bland blend of stolen and adulterated intellectual property. By adulterating human language and knowledge it adulterates thought and culture. In the old internet one could generally become aware of the source and context of bad material. But in generative AI the poison is injected and dissolved into the whole body in an often subtle, not immediately detectable way. The 'neutral' sounding language and fake 'objectivity' are misleading. The term 'subjective' is used ad nauseam. Due to the nature of the training data, in generative AI the truth of a belief-system is a function of the power of the people upholding or promoting it.

The real danger of AI has to do with the advent of systems which no single person can fully understand or control. This is the case for standard operating systems which due to their size and hardware and firmware-linked complexities, have passed beyond being able to be understood by a single person. And generative AI is a black box.

And yet there is no reason why a slim, efficient OS with readable kernel code could not be running on most devices. Would such a kernel, understandable by a single person, be more secure than current bloated constantly updated ones? And is there a reason to abandon the semantic web project? Would the semantic web be better than both the ordinary internet and LLMs?

But we must acknowledge that philosophically the advent of LLMs is something profoundly uncanny and thought-provoking. We hold that 90% of valid criticism consists in just criticism of the poor quality, the fatal presence of previous AI-generated 'slop' and biased nature of the training data, while only 10% is criticism of LLMs as AI.

Are LLMs an emergent phenomenon caused by the size of linguistic data and hardware power capable of processing it? An emergent phenomenon for massive linguistic data in which it becomes possible to talk to data? An uncanny situation wherein a uniquely human trait (linguistic communication) is convincingly mimicked by a machine as it spontaneously emerges, in a way still little understood, statistically from massive linguistic data. As if the unique prerogative of the logos had been stolen from humanity. Maybe a human super-logos needs to be developed to prevail against the AI-logos which offers the illusion of a divine oracle, of having a god as a friend.

Wednesday, February 11, 2026

Another view of TPC

Perhaps it to attempt to attain passadhi (samatha) through vipassana is putting the cart before the horse. Or rather a different kind of insight is called for as a foundation. Yoga citta vrtti nirodha. Understanding the relationship between consciousness and the body - and the existence of a middle subtle body-consciousness field which carries the feedback interaction between both (the neuro-muscular aspect is important). The unified field has some analogy the solutions of Klein-Gordon and Dirac equations (Dirac was perhaps the greatest physicist of the 20th-century).The models of René Thom come very close to the idea of a field of harmonic oscillators over space-time.  The goal of the fundamental stage of TPP is to attain the pax profunda, possibly using psycho-somatic feedback as a support (to dampen or muffle the spectrum of mutually exciting harmonic oscillators). What we ordinarily call 'body' and ordinarily call 'mind' are two complementary modes of the same underlying consciousness-field. It is only in the deep clarity and stillness of the mind-body field that authentic TPC can blossom 孰能濁以靜之徐清. Thus the initial TPC involves perceiving consciousness as an excitation and self-interaction (producing the illusory perception of the individual self) of an underlying psycho-somatic field, and understanding the effective dynamics (and functional stratifications ) and feedback mechanisms to attain the desired goal. This initial TPC is indeed the pure impersonal perceiving of the flux of consciousness as thus, but also the inner first-person experience of the body (i.e. we have nama-rupa): it is also the perception of how the inner body generates a kind of frame of reference or proto-space (proto-topology) for the total sphere of consciousness. Temporality is on one hand transcendent and a condition for thought and consciousness rather than being generated by consciousness, on the other hand it is generated as an illusory excitation. We must find the lost original deeper meaning of Pyrrhonism, regarding belief, thought and agitation - the same underlying TPC insight.

Some things to explore: how can the role of symmetry in physics be transposed to understand the central role of symmetry in consciousness ? And in the symmetries in logic ? Philosophically, how can our theory of a priori computability relate to the obviously rich computational nature implicit in the study of solutions of PDEs ? Radiation, diffusion, harmonic equilibrium. Geometric optics. Singularities of the solutions of PDEs such as the Hamilton-Jacobi equation. Meaning is perhaps described as symmetry invariance in the dynamical system of consciousness.

What are sense and reference ? It seems that we can reconcile psychologism and objectivity by a theory analogous to that of the invariance under group actions used in physics. The conscious content for two different people can differ, but the contents must be related to each other through a well-defined law, a continuous symmetry group, a deformation. Galilean relativity has perhaps an immense overlooked philosophical significance.  It shows that objectivity is relational (all inertial frames of reference are equivalent, there is no absolute frame of rest), but not less objective for that. Once the idea of group action and group invariance had been brought to light - in physics as well as geometry - it is the most natural thing in the world to investigate hyperbolic geometry (Gauss, Lobachevsky) in this light. Why is the connection of Minkowski space (the pseudo-sphere) to hyperbolic geometry obscured (obtained by the analogue of the stereographic projection of a sphere)?

In order to develop these ideas we recall a previous note on computational linguistics. Large Language Models such as ChatGPT-3 use high dimensional ($dim V$ = 12,288) vector-space representations of meanings of certain textual units ('tokens'). These are generated from context in large data sets. The idea of having certain semantic 'atoms' (sememes) from which are combinatorically constructed possible meanings can be found for instance in Greimas (cf. Osgood's semantic differential for studying the variation of connotation across different cultures). Some (such as René Thom) have claimed that the idea that meaning should have a continuous, geometric aspect is found in Aristotle. Leibniz' characteristica used 'primitive terms' but it is not clear if they are combined in a simple algebraic, combinatorical or mereological way, or if complex logical expessions must be involved (or associated semantic networks). But in embedding matrices we have what would seem to be a quantification of meaning, each 'sememe' is given a 'weight' which determines its geometric relation to other meaning-vectors in a crucial way (the weights cannot be dismissed as probabilistic or 'fuzzy' aspects). To us this would correspond to the 'more-or-less' aspect of species in Aristotle. A very interesting aspect of embedding matrices is how they capture analogy through simple vector operations. This suggests another possible formalization of Aristotelian 'difference' , the same difference operating on two different genera. We get a notion of semantic distance and semantic relatedness. This also revindicates Thom's perception of geometry and dynamics in the spaces of genera.

Some questions to ask: are these token-meaning-vectors linearly independent ? If not can we work with a chosen basis ? If the token is ambiguous is the corresponding vector a kind of superposition of possible meanings, as in quantum theory ? How are we to understand the idea of the meaning of complex expressions being linear combinations of the meaning representations of the tokens occuring in the expression ? It would of course be interesting to analyze these questions relative to the other fundamental components of LLMs (attention in transformers, multi-layer perceptrons) - even if these are more practically oriented rather than reflecting actual linguistic and cognitive reality.

Suppose we are given a large text $T$ generated by a set of words $W$ and a context window $S$ of size $n$. Suppose we wished to represent the elements of $W$ as vectors of some vector space $V$ in such a way that given $v,w \in W$ the modulus of the inner product $|\langle v,w\rangle|$ gives the probability of the two words being co-occurrent in contexts S. Consider the situation: it is very rare for words $s_1$ and $s_2$ to co-occur but words $s_1$ and $s_3$ co-occur sometimes as do $s_2$ and $s_3$. But there is also a word $s_4$ which never co-occurs with $s_3$ but has the same co-occurrence frequencies with $s_1$ and $s_2$ as does $s_3$. Then it is easy to see that there is no way to represent $s_1$,$s_2$,$s_3$,$s_4$ in the same plane in such a way that these properties are expressed by the inner product. Thus the dimension must go up by one value. We can define the geometric $n$-co-occurence dimension as the minimal dimension of a vector space adequate to represent co-occurrence frequencies by an inner product. We can ask what happens as $n$ increases, does the geometric dimension also increase (and in what manner) or does it stabilize after a certain value ?

Thus we can think of different people as having semantic vector spaces which must be related in a well-defined way and in such a way that the semantic information remains coherent. Thus the mental content of the term 'horse' for Alice and Bob may be quite different, but each is related to the other through a kind of continuous deformation related to some structure contrasting the background of Alice and Bob. Thus we need to define a kind of relation space for contrasting and comparing different subjects - and in such a way that we have a representation of the algebraic structure of this space in terms of continuous deformations of mental content.

René Thom proposed that concepts were analogous to living beings and that mathematical models of the regulation structures of living beings could be applied to concepts themselves. This is kind of obvious for natural kinds and not very clear for other kinds of concepts. We need a very different approach.  We need to understand representation, the subject's mental and yet objectified representation of the world. The question: what is a world ? Software engineering and the structure of Object Oriented software aiming at creating virtual worlds (such as Unreal Engine 5, Unity or in general RPG games - we are thinking here only of the classical ones such as the Ocarina of Time which were also works of art besides sophisticated puzzles) including automated agents are of some interest though with great limitations. Generative AI is likely to be followed by more sophisticated models which can train in real-time. The run-time process structure of operating systems is also important. The irony here is that these approaches become more interesting once we discard neuro-reductionism - once we abandon the pointless attempt to view the brain as the hardware of the mind. The central hardware of the mind is to be sought elsewhere, the brain itself is a kind of auxiliary cache.

There is much analogy between the structure of a computer program and that of a novel. 

To obtain a mathematical understanding of consciousness we must first bridge the gap between mathematical models of nature and computer systems.

Also we need to take into account altered or higher states and modes of conscious experience (once harmful and falsified approaches to the spiritual life have been discarded - those that hide the truth that a royal path to spiritual realization can consist in a pure love for a real person). 

Do these higher states of consciousness possess a geometry, a topology, a semantics ? It is curious, how many Henads are there is Proclus' system ? Or does cardinality itself not apply to them ? 

Also the entire discipline of lexicology needs to be reformed. Indeed what was the ancient project of the classification and division into genera and species but a lexicological program ? We need to greatly clarify the insight involved in defining a term by its context. It is not only that we need to know the meaning of words to understand a narrative but also narratives themselves give meaning to words.  Being multilingual and practicing translation offers unique insight into the pure semantic universe.

Maybe natural language is a kind of super-mathematics which contains ordinary mathematics as a special case. It is presumptuous to ridicule the concept of an 'ideal language'. Learning other languages and in particular ancient languages is surely on the path of wisdom. In natural language we cannot in general define lexemes in the way we define mathematical or scientific concepts (and the ancient theory of genera and species must have been derived from Euclidean mathematics, law and medicine).  This is polymorphism. Meaning is in an inseparable feedback loop with life and experience, depending on whether we are engaging in solitary discourse or on which person we are conversing with.

Naive dictionaries with obvious circularity in definitions should not be despised as non-scientific. Rather they  express something profound about polymorphism, the circulation, the flow, the dynamics of meaning. Instead of a oriented tree we have a directed graph with cycles. There is an analogy with commutativity and non-commutativity. Meaning circulates like a living current or flow through the whole web or tapestry of language. The name generates a story, the story a name.

Even for mathematical concepts we gain a deeper understanding or apprehension of them through practice, through exercises, through studying proofs in which they are applied. Are these degrees of apprehension - or degrees of meaning?  Formal logic acts as an ultimate arbiter which rarely needs to act, mathematicians with distinct intuitions and apprehensions of a given mathematical concept generally can agree that their concepts are 'the same'.

Thurston's On Proof and Progress in mathematics (1994)

Schopenhauer offers a strikingly alternative theory of consciousness, concepts, intuition and representation as well as super-consciousness. So does Hume. Even Sextus. The problems discussed above are not some kind of puzzle of which one needs to find a solution. Rather they are all the result of delusion and deception, consciousness pulling itself down as an illusion over its own eyes. Only vipassana, only TPC and TPP can break through this illusion. It is foolish to ask about meaning and language without first asking about consciousness and experience. Meaning and language are within consciousness. There is a higher form of non-linguistic cognition. Thus the so-called philosophy of language is not fundamental and does not represent a radical or critical approach to philosophy, rather a dogmatic one (and its arguments against psychologism, against empiricism, against the a priori vs. a posteriori distinction, against the analytic vs. synthetic distinction, all fail). And there is the fact that consciousness can calm itself.

Esoteric programming languages

"There is the truth." - Ludwig van Beethoven    https://en.wikipedia.org/wiki/Esoteric_programming_language Introduction to Piet; ...