Showing posts with label humanity. Show all posts
Showing posts with label humanity. Show all posts

Saturday, July 4, 2026

A thesis on generative AI

We present to following thesis on generative AI.  That to understand fully the essence of generative AI we must first understand the socio-cultural and philosophical essence of the advent of the internet.  Generative AI is only possible not merely because of raw hardware capabilities and certain neural net architectures and training processes, but on a more fundamental level because of the nature of the large data sets employed. Generative AI is only possible because of the previous availability of a very special kind of massive data set - precisely the kind of data sets provided by the internet. Generative AI is simply the natural further development of the same factors at work in the genesis of the internet.  So our main task is to describe the nature of internet data as well as its social, psychological, linguistic and cultural correlates.  This is a difficult task. One should consider and evaluate the following aspects: the democratization of knowledge and expertise (what certain authors once called "the cult of amateurism"), the lessening of the both the physical and semantic volume of content - flattening of layers of semantic depth, connotation, subtext and extra-textual reference, a  massive atomization and repetition with small variants of certain templates, memes and "virality", propagation and reproduction of persistent psycho-linguistic patterns throughout the whole body of the data - all which structurally contribute to forming the kind of atomized, grid-like, decentralized, self-entangled and self-connected  massive data set which is precisely what allows the generative AI technology to produce LLMs which perform as they do. While traditional cultural-linguistic data involved relatively few authors with massive linguistic content, the internet data involves a massive amount of authors with relatively scarce linguistic content each.  Humanity is caught itself in the textual universe of the internet.  The internet data set can be described as an enmeshed chaos of micro cognitive-linguistic units - which we can compare to the set P of partial information used in forcing. If we wish to combat generative AI we need to get at the root of the whole historical-cultural process involved in the genesis of the internet. We need to instigate a revolution regarding authorship, epistemic authority and the nature of the written text. The semantic web project should be revived. Curiously enough the Pali canon is an very early example of large textual data which exhibits some of the characteristics of internet data.

If the internet era was the era of communication, personal development through sharing and community, but now not only has the "internet" as a medium become content itself but it has become dead content, i.e., generative AI. Thus we are in the heroic age in which the human person must again find development and peace in isolation again - so that again there is something to share. The highest form of human knowledge must be global and synthetic, partial analysis has as its ultimate aim furnishing the support for such a knowledge. Global and synthetic knowledge is the only knowledge that frees one from Plato's Cave. 

Can we reverse engineering generative AI using generative AI? Could we train an AI model M on the data D = (T,L) consisting of training data T and a LLM L trained on T, to solve the following problem:

Given a prompt P fed into L and response R, identify the portions of T which can be considered as having (in some mode or another) the largest weight in the generation of R from P.

Can we define the concept of essential dependence of response R to P via L on a portion S of T? Could we train M to return, for a given P, portions of T which have a certain structural similarity to R?  

Monday, April 27, 2026

On generative AI

Is generative AI corrupting human knowledge and language and by extension human thinking and human culture themselves?

A wikipedia dump is around 100 GB. Wikipedia could be improved and be semantically formatted to be computer readable and advanced query systems could be developed. Would not this be better for the acquisition of knowledge and the advancement of science? Are AI generated summaries of books or papers valid replacements for human ones? What justifies our trust in generative AI as compared to a search engine?

Generative AI is corrupting the internet. Maybe it is a zombie or Frankenstein of human language and knowledge. Or a bland blend of stolen and adulterated intellectual property. By adulterating human language and knowledge it adulterates thought and culture. In the old internet one could generally become aware of the source and context of bad material. But in generative AI the poison is injected and dissolved into the whole body in an often subtle, not immediately detectable way. The 'neutral' sounding language and fake 'objectivity' are misleading. The term 'subjective' is used ad nauseam. Due to the nature of the training data, in generative AI the truth of a belief-system is a function of the power of the people upholding or promoting it.

The real danger of AI has to do with the advent of systems which no single person can fully understand or control. This is the case for standard operating systems which due to their size and hardware and firmware-linked complexities, have passed beyond being able to be understood by a single person. And generative AI is a black box.

And yet there is no reason why a slim, efficient OS with readable kernel code could not be running on most devices. Would such a kernel, understandable by a single person, be more secure than current bloated constantly updated ones? And is there a reason to abandon the semantic web project? Would the semantic web be better than both the ordinary internet and LLMs?

But we must acknowledge that philosophically the advent of LLMs is something profoundly uncanny and thought-provoking. We hold that 90% of valid criticism consists in just criticism of the poor quality, the fatal presence of previous AI-generated 'slop' and biased nature of the training data, while only 10% is criticism of LLMs as AI.

Are LLMs an emergent phenomenon caused by the size of linguistic data and hardware power capable of processing it? An emergent phenomenon for massive linguistic data in which it becomes possible to talk to data? An uncanny situation wherein a uniquely human trait (linguistic communication) is convincingly mimicked by a machine as it spontaneously emerges, in a way still little understood, statistically from massive linguistic data. As if the unique prerogative of the logos had been stolen from humanity. Maybe a human super-logos needs to be developed to prevail against the AI-logos which offers the illusion of a divine oracle, of having a god as a friend.

Theory of Meaning

Meaning, that most illusive of philosophical concepts, is without doubt a ternary relation M(A,B,C): A means B relative to/in/according to C...