Showing posts with label machine learning. Show all posts
Showing posts with label machine learning. Show all posts

Friday, July 31, 2026

Mathematics and AI

Referring to certain software and hardware as "AI" and to such alleged "AI" as an agent "solving" problems in mathematics is profoundly misleading. The programs in question are not "intelligent" and are not "solving" mathematical problems as agents in any philosophically meaningful sense. And these programs will certainly never "replace" human scientific endeavor in any conceivable way, rather they are, and will ever remain, mere tools.

Let us be clear. Human beings can only process, check and produce data within definite finite bounds based on symbol systems with definite finite bounds and bounded fragments of finitarily determined rules.

Thus it is a triviality and a truism that all human external symbolic activity and productions could in principle - given a massive enough set of data is made available - be mimicked and processed by brute-force. We can imagine a supercomputer in space with processing power and storage a billion times surpassing any of the human bounds of symbol processing and text processing, checking and production. We can also give our supercomputer some kind of super-luminal processing velocity. This supercomputer through brute-force and crude machine learning algorithms would beat and outperform every man-made "AI" in every possible domain in an instant. There is nothing surprising here and there is nothing here that has anything remotely to do with "intelligence".

Intelligence is rather reflected in doing much with little external, material, processing power. Chess programs are just cheating machines. Our hypothetical supercomputer would beat any current Go or Chess AI and that would not make it "intelligent" in any meaningful way.

As for mathematics, let us take a proof assistant such as Agda or Idris 2. The mathematician develops a theory (or formalizes a previous theory) by means of type definitions, records and type declarations for terms. The problem is to find explicit proof terms. This could be attempted by using brute-force, heuristics or any kind of machine learning method. And this search could fail. Or the search algorithm could be refined and altered. There is nothing unusual going on here. The algorithm is not "doing mathematics", it is not constructing theories, refining definitions, improving its own heuristics or being telologically oriented towards a certain architectural vision. It is the mathematician's legitimate tool.

Speaking of "AI" "solving" mathematical problems or "doing" mathematics or potentially "replacing" mathematicians is sheer and utter nonsense.

If even published journal papers can contain errors, I certainly would not trust any "mathematical" output produced by generative AI that was not formalized and checked in all its details by current proof assistants and proof checkers.

Saturday, May 9, 2026

Short philosophical considerations on AI

Hegel and Heidegger were thinkers about their own time, thinkers about historical events and happenings. Few have the insight and courage to fathom the full depth of the meaning of an historical event, the coming to be (coming of age?) of an historical process. Tragically,  it is only some time after the event (the time of monsters?) has hit humanity with full force that Hegel's famous owl can spread her wings. Is it not true that some of what the prophetic author of Sein und Zeit wrote about technology only makes full sense at the present?

This seems to us particularly true of the emergence of the age of the Internet and the age of generative AI which is its logical development.

And yet nothing could be further from our own philosophizing than any form of historicism or historical philosophy. As such both Hegel and Heidegger, for all their interest and insight, must be considered as having crafted systems based on an incorrigible error. 

Social progress is not a law of nature but a legitimate hope - even if at present it seems a distant one - and it is our moral duty to work towards it in the midst of uncertain outcomes.

The advent of the internet was the advent of connection between people. This connection carried rhetorically moral undertones and echoed enlightenment ideals about the desirability of sharing and making knowledge available.  In the present age of the generative AI based Internet powered and controlled by corporations and governments aligned to anti-enlightenment ideals, it may be that it is morally called upon us to practice instead the process of disconnection and the purification and preservation of knowledge(not obviously in the sense of the 'great simplification' of the Canticle of Leibowitz).   

The most basic step is ensuring locality of core information. That is, to be in possession of machines onto which have been downloaded significant portions of Internet Encyclopedias (despite their serious shortcomings) as well as some decently performing LLM. To this we add, it needs not be said, massive of digital preservation of human cultural artifacts, notably libraries.

One can use Kiwix and download for offline viewing the most recent English Wikipedia (50GB text-only 150GB with pictures).  With a AMD Ryzen 7 5825U processor with 16GB RAM and 2GB Radeon Graphics one can use Ollama and download and use some decently performing LLMs (gemma4 comes in E2B, E4B, 31B and 26B A4B). 

Most living beings alternative between states of being awake and of sleep. Can would we design a dynamic LLM which similarly alternates between states of user interaction and re-training based on this interaction? The most important being the correction and/or updating of knowledge or perhaps the removal of harmful and biased content and "thought patterns". If LLMs can improve then it is not only  a question of having the number of parameters equal to the number of neurons of the human brain.

Computers can enhance and aid human cognition as well as hinder and destroy it (there is a growing body of evidence concerning the disastrous effect of excessive or inappropriate generative AI use for individual mental health and cognitive development, not to mention for society as a whole). 

But the harms of generative AI have little to do with lesser-known extremely powerful and beneficial aspects of the computer for human cognition. We cannot go into this in detail here. Let it just be said that it involves using adequate software for the rigorous formalization of human scientific theories and concepts (specially logic and mathematics and formal methods in the sciences) and the vital feedback-loop between human thought and the software interface (IDE) which results in the simultaneous enhancement of human understanding and production and the quality of the software-based formalization and implementation. 

The software in question includes not only Rocq (formerly Coq), Agda and functional programming language but such languages as Python, Javascript and C/C++. Python is a multi-paradigm and highly versatile language with an elegant syntax. Python comes close to achieving the ideal of a universal language in the Leibnizian sense and is a wonderful tool for formalization, implementation,  verification and exploration in a variety of areas in mathematical logic and finite mathematics. 

We note also the importance of minimalism (using as few dependencies as possible) and building things from the ground up - this goes for scientific and philosophical projects, not of course for commercial and industrial ones.  We will address in the future the question of the possible role of machine learning in this process.

Monday, April 27, 2026

On generative AI

Is generative AI corrupting human knowledge and language and by extension human thinking and human culture themselves?

A wikipedia dump is around 100 GB. Wikipedia could be improved and be semantically formatted to be computer readable and advanced query systems could be developed. Would not this be better for the acquisition of knowledge and the advancement of science? Are AI generated summaries of books or papers valid replacements for human ones? What justifies our trust in generative AI as compared to a search engine?

Generative AI is corrupting the internet. Maybe it is a zombie or Frankenstein of human language and knowledge. Or a bland blend of stolen and adulterated intellectual property. By adulterating human language and knowledge it adulterates thought and culture. In the old internet one could generally become aware of the source and context of bad material. But in generative AI the poison is injected and dissolved into the whole body in an often subtle, not immediately detectable way. The 'neutral' sounding language and fake 'objectivity' are misleading. The term 'subjective' is used ad nauseam. Due to the nature of the training data, in generative AI the truth of a belief-system is a function of the power of the people upholding or promoting it.

The real danger of AI has to do with the advent of systems which no single person can fully understand or control. This is the case for standard operating systems which due to their size and hardware and firmware-linked complexities, have passed beyond being able to be understood by a single person. And generative AI is a black box.

And yet there is no reason why a slim, efficient OS with readable kernel code could not be running on most devices. Would such a kernel, understandable by a single person, be more secure than current bloated constantly updated ones? And is there a reason to abandon the semantic web project? Would the semantic web be better than both the ordinary internet and LLMs?

But we must acknowledge that philosophically the advent of LLMs is something profoundly uncanny and thought-provoking. We hold that 90% of valid criticism consists in just criticism of the poor quality, the fatal presence of previous AI-generated 'slop' and biased nature of the training data, while only 10% is criticism of LLMs as AI.

Are LLMs an emergent phenomenon caused by the size of linguistic data and hardware power capable of processing it? An emergent phenomenon for massive linguistic data in which it becomes possible to talk to data? An uncanny situation wherein a uniquely human trait (linguistic communication) is convincingly mimicked by a machine as it spontaneously emerges, in a way still little understood, statistically from massive linguistic data. As if the unique prerogative of the logos had been stolen from humanity. Maybe a human super-logos needs to be developed to prevail against the AI-logos which offers the illusion of a divine oracle, of having a god as a friend.

Friday, November 14, 2025

Attention is all you need

Around 2017 I was thinking about the problem of automatic translation and how ambiguity and idioms could be dealt with. One idea was to tag every word $w$ with one of Roget's 1000 categories (represented by a set $C$). Thus we have map $\kappa: W \rightarrow P(C)$.  Roughly speaking an ambiguous word $w$ will have $\kappa(w)$ with cardinality at least 2. Given a context $T$ and a word $w$ occurring in $T$  my idea was to devise an algorithm which functioned a little bit like a Sudoku puzzle using a concept of 'semantic distance'.  We find a word $w$ such that based on the current words $v$ in the context with singleton $\kappa$ we can determine which $c \in \kappa(w)$ is 'closest' to the set of $c$'s inhabiting the singletons of such words. We then make the choice and this should lead to finding further words that can be resolved and so forth. Of course the problem is how to define such semantic distance as well as to guarantee that the process achieves its goal and does not get stuck (but we could introduce random choices). If we view Roget's 1000 categories as organized as leaves (or even nodes) of a binary tree then there is an obvious definition. For instance 'rotation' is semantically closed to 'motion' than it is to 'feeling'.

There is a problem with compositionality for idioms like 'raining cats and dogs' or for a term like 'white rhinocerous'. Thus composition of meanings is in general multi-valued. We have uploaded a small text about this on researchgate and other platforms.

The vector representations used in LLMs suggest the following speculation. Could it be that meaning can be coherently constructed out of complex entities which themselves have no intrinsic (or easily assignable) meaning ? Or like in quantum mechanics the wave function (proto-meaning) is essentially a superimposition of eigenstates corresponding to actual observables (meanings). 

Around 2017 I was thinking about the problem of automatic translation and how ambiguity and idioms could be dealt with. One idea was to tag every word $w$ with one of Roget's 1000 categories (represented by a set $C$). Thus we have map $\kappa: W \rightarrow P(C)$. Roughly speaking an ambiguous word $w$ will have $\kappa(w)$ with cardinality at least 2. Given a context $T$ and a word $w$ occurring in $T$ my idea was to devise an algorithm which functioned a little bit like a Sudoku puzzle using a concept of 'semantic distance'. We find a word $w$ such that based on the current words $v$ in the context with singleton $\kappa$ we can determine which $c \in \kappa(w)$ is 'closest' to the set of $c$'s inhabiting the singletons of such words. We then make the choice and this should lead to finding further words that can be resolved and so forth. Of course the problem is how to define such semantic distance as well as to guarantee that the process achieves its goal and does not get stuck (but we could introduce random choices). If we view Roget's 1000 categories as organized as leaves (or even nodes) of a binary tree then there is an obvious definition. For instance 'rotation' is semantically closed to 'motion' than it is to 'feeling'. There is a problem with compositionality for idioms like 'raining cats and dogs' or for a term like 'white rhinocerous'. Thus composition of meanings is in general multi-valued. We have uploaded a small text about this on researchgate and other platforms. The vector representations used in LLMs suggest the following speculation. Could it be that meaning can be coherently constructed out of complex entities which themselves have no intrinsic (or easily assignable) meaning ? Or like in quantum mechanics the wave function (proto-meaning) is essentially a superimposition of eigenstates corresponding to actual observables (meanings). In this note we present an abstract formal approach to the basic problems regarding texts and meaning.

We start with a non-empty set $W$ of word-forms, expressions which have no meaning-bearing parts. We consider a subset $T \subset W^\star$ of possible meaningful texts. $T$ must satsify the following conditions: \[ W \subset T\] \[ \forall t \in T \, t\notin W \rightarrow \exists s,u \in T \, t = su\] The last axiom means that every text has at least on syntactic decomposition - a binary tree expressing sucessive division into meaningful elements down to the level of $W$. 
We are given furthermore a non-empty set $M$ of possible meanings. We postulate a map \[\sigma : T \rightarrow P(M)\] satisfying $\sigma(t) \neq \emptyset$ for all $t \in T$ and a map \[\lambda : M \rightarrow P(T)\] satisfying $\lambda(m) \neq \emptyset$ for all $m \in M$. The first map gives the possible meanings of a text $t$ and the second map gives the possible ways to express a given meaning $m$.
Recall that for any set we have a map $P(P(X)) \rightarrow P(X)$ obtained by taking unions. Using this map we can compose $\sigma$ and $\lambda$ to obtain maps \[\pi : T \rightarrow P(T)\] and \[\alpha : M \rightarrow P(M)\] which express the ways to paraphrase a given text and the possible linguistic ambiguities (semantic misreadings) conveyed by expressions of a given meaning.

Given $s,t \in T$ we define $s \subset t$ to mean that there exists $s_1,...,s_n \in T$ and $u_2,...,u_{n-1} \in T$ such that (for $ n> 1$, $s_n = t$,  $s_i = s_{i-1}u_{n-1}$ or $s_i = u_{n_1}s$ for $i > 1$ and $s_1 = s$. For $n = 1$ we require that $s = t$.  We postulate that there is a family of maps $\rho_{ts} : \sigma(t) \rightarrow \sigma(s)$ for each $t,s \in T$ with $s \subset t$ which satisfies the properties $\rho_{tt} = id$ and $\rho_{su}\circ\rho_{ts} = \rho_{tu}$. Thus we boldly state that there is a restriction to 'white' of the standard meaning of 'white rhinocerous'. This means that the meaning of a text uniquely determines the meanings of all its meaningful components.

Furthermore $\rho$ must satisfy the fundamental property: Given $t = s_1s_2$ with $s_1,s_2 \in T$ and given $m_1 \in \sigma(s_1)$ and $m_2 \in \sigma(s_2)$ there exists at least one $m \in \sigma (t)$ such that $\rho_{ts_1}(m) = m_1$ and $\rho_{ts_2}(m) = m_2$. An important corollary is that if $s \subset t$ then for every $m \in \sigma(s)$ there exists $n \in \sigma(t)$ such that $\rho_{ts}(n) = m$. We hold that a lawful syntactical combination of meaningful expressions has some meaning. Indeed 'green idea' is just as meaningful as Polonius' 'green girl'. Metaphor should be at the heart of any cogent linguistic philosophy and formal linguistics. But our principle might seem to fail here because the restriction of the meaning of this expression to 'green' has a different meaning from the ordinary perceptive meaning of 'green'. Metaphoric green is a valid meaning for the expression 'green' and thus the meaning of the combination 'green idea' is valid. However even the expression 'green idea' with 'green' in its perceptual sense is just as meaningful as Meinong's 'square circle'.

The new senses of 'green' and 'white' adquired by restriction from 'green girl' and 'white rhinocerous' recall some of the mechanisms for attention in LLMs. We define the set of $t$-contexts for $t\in T$ to be $C(t) = \{(t_1,t_2) : t_1,t_2 \in T \& t_1t_2 \in T\}$. For $c = (t_1,t_2) \in C(t)$ we write $c[t]$ for $t_1tt_2$.

Finally we postulate that, for $t \in T$, there is a map $\kappa : \sigma(t) \rightarrow C(t)$ satisfying the following property for all $m \in \sigma(t)$: \[ \forall n \in \sigma (\kappa (m)[t])\, \rho_{\kappa(m)[t] t}(n) = m \]  This means that given an expression $t$ there is a context such that even if ambiguous still determines the meaning of $t$ uniquely. We can think of contexts for 'white rhinocerous' involving 'stuffed animal' and involving 'species' which determine distinct meanings.

The above theory could be carried out using the set of terms generated freely by a non-associative operation $W^\bullet$ and by postulating that each $t \in T$ has a unique syntactic decomposition tree in $W^\bullet$ - but this is a little too restrictive perhaps. We can rewrite the above theory in probabilistic terms, specially $\kappa$. We can replace contexts with co-occurrences. And turn $\kappa$ into map which from the data of certain $t_i$s co-occuring with $t$ (at certain distances) assigns a certain probability distribution to possible meanings of $t$.

But let us now consider another approach. We suppose that there is a 'metric' defined on $M$, $d: M \times M \rightarrow \mathbb{R}_0^+$. To simplify things we can also consider a finite disjoint decomposition of $M$ into categories $M = \bigcup_{c\in C} M_c$ (for example something based on the 1000 categories of Roget' thesaurus, but these of course only work for $W$ or small texts) and consider the distance only on $C$ and extend it in a trivial way to $M$. And important question: what is the distance of $m \in \sigma(ts)$ to $\rho_{(ts)t}(m) \in \sigma(t)$ ? Note that each $t \in T$ determines a subset $\sigma(t) \subset M$. Suppose we have a context $c \in C(t)$ and consider $\sigma(c) \subset M$. How do we define the distance between a $x \in M$ and a subset $ X \subset M$ ? One possibility is $d(x,X) = min \{d(x,z) : z \in X \}$ but this is not what we want. Rather we need an average over all the distances $d(x,z)$ for $z \in X$. For finite $M$ this is easy. Otherwise we need to introduce a measure on $M$. Now we can consider a map  \[\kappa: C(t) \rightarrow \sigma(t) \]  which associates to each $c \in C(t)$ the element $m \in \sigma(t)$ such that $d(m, \sigma(c))$ is smallest. Or from a practical perspective we can use the distance induced by the finite decomposition into categories. This works best (for the case of Roget's categories) if instead of using $\sigma(c)$ we use the subset of $M$ determined by all of the $\sigma(w)$ for $w$ occurring in $c$. The problem with this approach is that the minimum may not be unique. But for a text $t$ we can decompose it into a context $c[s]$ for different $s \subset t$. Our algorithm would first seek the right $s$ so that $\kappa$ determines a unique minimum. This will allows us to refined the images in $M$ for the other contexts (strickly speaking this is no longer the union of the $\sigma(w)$ but a refined subset). It is to be hoped that this process could be continued until the entire $t$ is disambuiguated. Is this not what the human mind does ?  It may be worth considering that $M$ has a more refined description in terms of subsets of a certain 'semantic space' $\Sigma$ in such a way that the $\sigma$ of elements of $W$ are like 'points', their concatenations like 'lines' and so forth. And we can define the metric on $\Sigma$ rather than directly on $M$.  Note that we must not confuse the mathematical with the philosophical aspect. If we postulate that meaning can be formally analized and studied via context this is not meant to imply that meaning is context, anymore than studying a recursive set as a set means that we are identifying the associated algorithm with the set itself.

Given a $T \subset W^\star$ (for simplicity we assume it is finite) as in section 1, let us consider a map $c: T \rightarrow P(T)$ which associates to each $t \in T$ the set set of all $s' \in T$ such that $t$ occurs in $s$ (contexts). From a practical perspective it is better to restrict $c$ to contexts having a certain limit of length, which we fix to be $n$ (and denote by $T_n$) and denote the resulting map by $c_n$. Of course then $c_n$ becomes empy for $t$ precisely of length $n$. We can define a notion of distance $d(t_1,t_2)$ between $t_1,t_2 \in T$. This is given by  \[\frac{\Sigma_{t \in T_n: t_1,t_2 \in t} dist(t_1,t_2,t) }{|T_n|}\]  where $dist$ is the ordinary distance between $t_1,t_2$ in $t$. There are various ways to define distances of sets. In our case we should use $d(A,B)$ being equal to the average distance. The idea is to decompose $c_n(t)$ into subsets which are far apart, reflecting distinct semantic categories. If for a given radius the elements of $T_n$ in that radius are random and dispersed then this is not possible. Ideally we wish to find 'clusters'. We can view the decomposition into clusters as writing a vector represeting a general element as a linear combination of vecotrs representing the clusters or categories.

Note that if $t_1t_2$ is know to be part of a defined cluster then it is plausible that this cluster also determines one for $t_1$ and $t_2$ thus revindicating the existence of the map $\rho$. Let us look more closely at $T$. It can mean all possible texts in a given language.  But if we consider a person $p$ then we can take $T_p$ to mean the subset generated by the meaningful subsegments of the text $t_p$ consisting of all linguistic material either thought in inner verbal discourse, heard or spoken from the moment of birth to the moment of death. Or we can consider in some sense all the possible such maximal onto-texts. We can also define the same kind of text for communities and their history. $T$ is then like a book whose pages correspond to a possible total text of linguistic material in a certain possible history of a community (cosmo-texts). The set of possible books of the world encompassing the total linguistic-consciousness of all people throughout history.  A text is mainly about certain topics. A biography is about a person. Can we give a purely textual definition of aboutness ?

Since a cosmo-text is finite there will be meaningful texts which will escape it. Is there an English text of 1000 words which will never be read by mankind and yet is highly meaningful (or would be meaningful in its possible encompassing cosmo-text) ? Can we give a purely textual combinatorial definition of 'highly meaningful' ? We give an abstract answer. Consider a set $\Sigma$ of symbols and suppose we have a metric $d$ on $\Sigma^\star$. Let $S \subset \Sigma^\star $. Let $S(w)$ denote all elements in $S$ with initial segment $w$. Then a word $w \in \Sigma^\star$ more impactful (meaningful) than $w' \in \Sigma^\star$ for $s$ in $S$ if the rough 'distance' between $S(sw)$ and $S(sw')$ is significant (again we need a good notion of distance between sets). This would correspond to an 'influential' or 'seminal' work in the global cosmo-text.

Large Language Models such as ChatGPT-3 use high dimensional ($dim V$ = 12,288) vector-space representations of meanings of certain textual units ('tokens'). These are generated from context in large data sets. The idea of having certain semantic 'atoms' (sememes) from which are combinatorically constructed possible meanings can be found for instance in Greimas (cf. Osgood's semantic differential for studying the variation of connotation across different cultures). Some (such as René Thom) have claimed that the idea that meaning should have a continuous, geometric aspect is found in Aristotle. Leibniz' characteristica used 'primitive terms' but it is not clear if they are combined in a simple algebraic, combinatorical or mereological way, or if complex logical expessions must be involved (or associated semantic networks). But in embedding matrices we have what would seem to be a quantification of meaning, each 'sememe' is given a 'weight' which determines its geometric relation to other meaning-vectors in a crucial way (the weights cannot be dismissed as probabilistic or 'fuzzy' aspects). To us this would correspond to the 'more-or-less' aspect of species in Aristotle. A very interesting aspect of embedding matrices is how they capture analogy through simple vector operations. This suggests another possible formalization of Aristotelian 'difference' , the same difference operating on two different genera. We get a notion of semantic distance and semantic relatedness. This also revindicates Thom's perception of geometry and dynamics in the spaces of genera. Some questions to ask: are these token-meaning-vectors linearly independent ? If not can we work with a chosen basis ? If the token is ambiguous is the corresponding vector a kind of superposition of possible meanings, as in quantum theory ? How are we to understand the idea of the meaning of complex expressions being linear combinations of the meaning representations of the tokens occuring in the expression ? It would of course be interesting to analyze these questions relative to the other fundamental components of LLMs (attention in transformers, multi-layer perceptrons) - even if these are more practically oriented rather than reflecting actual linguistic and cognitive reality.  Suppose we are given a large text $T$ generated by a set of words $W$ and a context window $S$ of size $n$. Suppose we wished to represent the elements of $W$ as vectors of some vector space $V$ in such a way that given v,w in W the modulus of the inner product $|\langle v,w\rangle|$ gives the probability of the two words being co-occurent in contexts S. Consider the situation: it is very rare for words $s_1$ and $s_2$ to co-occur but words $s_1$ and $s_3$ co-occur sometimes as do $s_2$ and $s_3$. But there is also a word $s_4$ which never co-occurs with $s_3$ but has the same co-occurrence frequencies with $s_1$ and $s_2$ as does $s_3$. Then it is easy to see that there is no way to represent $s_1$,$s_2$,$s_3$,$s_4$ in the same plane in such a way that these properties are expressed by the inner product. Thus the dimension must go up by one value. We can define the geometric $n$-co-occurence dimension as the minimal dimension of a vector space adequate to represent co-occurrence frequencies by an inner product. We can ask what happens as $n$ increases, does the geometric dimension also increase (and in what manner) or does it stabilize after a certain value ?

Svetla Slaveva-Griffin, Plotinus on Number (2009)

https://bmcr.brynmawr.edu/2010/2010.02.17/ Ennead VI,6, which deals with Plotinus' philosophy of number, is a very difficult text to und...