AI: Understanding It, Judging It, and Shaping What Comes Next
Core thesis
AI should be used to increase understanding, not to outsource understanding.
And it should be used to assist judgment, not to replace judgment.
That is the thread running through everything here: the mathematics, the economics, the ethics, the art, the environmental questions, and even the geopolitical questions.
AI is useful precisely because it can compress enormous amounts of information and surface patterns that would otherwise take us much longer to find. But if we use it merely to obtain an answer without understanding why the answer is true, what assumptions it depends on, what alternatives were discarded, or who bears the consequences if it is wrong, then we have not really gained intelligence. We have gained dependence.
The goal is not to become anti-AI. It is not to become pro-AI either.
The goal is to understand it well enough that we can use it without worshipping it, fear it without inventing things about it, regulate it without freezing it in place, and benefit from it without handing over our responsibility to think.
1. The history of AI and what it actually is
The easiest way to understand modern AI is not to start with ChatGPT.
Start with algebra.
At the simplest level, algebra asks us to find an unknown.
\[ 2x + 3 = 11 \]We solve for one variable:
\[ x = 4 \]That is already a kind of inference. We have known relationships, one unknown, and a rule for finding it.
Then we generalize.
Instead of one variable, suppose we have several:
\[ a_1x_1 + a_2x_2 + \cdots + a_nx_n = y \]Now we are in multivariable algebra.
When there are many such equations, it becomes much cleaner to write them as linear algebra:
\[ A\mathbf{x} = \mathbf{y} \]where:
- \(A\) is a matrix,
- \(\mathbf{x}\) is a vector of unknown values,
- \(\mathbf{y}\) is the result.
This is the conceptual doorway into neural networks.
1.1 A simple teaching example
Suppose we have two people:
- Alwyn
- Thomas
and two possible foods:
- ice cream
- pizza
We can encode the people with one-hot vectors:
\[ \text{Alwyn} = \begin{bmatrix} 1 & 0 \end{bmatrix} \]\[ \text{Thomas} = \begin{bmatrix} 0 & 1 \end{bmatrix} \]Then imagine a learned matrix:
\[ M = \begin{bmatrix} 0.2 & 0.8 \\ 0.7 & 0.3 \end{bmatrix} \]If the columns mean:
\[ [\text{ice cream},\ \text{pizza}] \]then:
\[ \begin{bmatrix} 1 & 0 \end{bmatrix} M = \begin{bmatrix} 0.2 & 0.8 \end{bmatrix} \]The model currently associates Alwyn more strongly with pizza.
Do the same for Thomas:
\[ \begin{bmatrix} 0 & 1 \end{bmatrix} M = \begin{bmatrix} 0.7 & 0.3 \end{bmatrix} \]Now Thomas is more strongly associated with ice cream.
This is deliberately tiny, but the basic idea scales surprisingly far.
A modern neural network is still doing repeated numerical transformations of vectors with learned parameters.
The scale is just absurdly larger.
1.2 From exact answers to least-squares error
Real data does not normally fit our equations perfectly.
Suppose we want:
\[ A\mathbf{x} \approx \mathbf{y} \]Then instead of asking for an exact solution, we can minimize the error.
A classical choice is least-squares error:
\[ L = \sum_i (y_i - \hat{y}_i)^2 \]or in vector form:
\[ L = \|\mathbf{y} - A\mathbf{x}\|_2^2 \]This is a major conceptual step.
We are no longer asking:
What is the exact answer?
We are asking:
What parameters make the answer as close as possible to what we observed?
That is already the core idea behind training.
1.3 The gradient
Now suppose there are many parameters.
We need to know which way to change them to reduce the loss.
For a function:
\[ L(w_1,w_2,\ldots,w_n) \]the gradient is:
\[ \nabla L = \begin{bmatrix} \frac{\partial L}{\partial w_1} \\ \frac{\partial L}{\partial w_2} \\ \vdots \\ \frac{\partial L}{\partial w_n} \end{bmatrix} \]The gradient points toward increasing loss.
So gradient descent goes the other way:
\[ \mathbf{w}_{t+1} = \mathbf{w}_t - \eta \nabla L \]where \(\eta\) is the learning rate.
That equation is one of the most important equations in modern machine learning.
The important correction is this:
We do not minimize the gradient.
We minimize the loss.
The gradient tells us which direction changes the loss.
1.4 The perceptron
One of the early neural models is the perceptron.
At its simplest:
\[ z = \mathbf{w}^T\mathbf{x} + b \]and then:
\[ y = f(z) \]where \(f\) is an activation function.
Originally, a simple threshold might be used:
\[ f(z) = \begin{cases} 1, & z > 0 \\ 0, & z \le 0 \end{cases} \]That gives us a linear decision boundary.
One perceptron can separate some problems.
It cannot separate everything.
The famous XOR problem helped make this limitation painfully clear.
So researchers added layers.
1.5 Multiple layers and changing topology
Once we stack transformations, we get something like:
\[ \mathbf{h}_1 = f(W_1\mathbf{x} + \mathbf{b}_1) \]\[ \mathbf{h}_2 = f(W_2\mathbf{h}_1 + \mathbf{b}_2) \]\[ \mathbf{y} = W_3\mathbf{h}_2 + \mathbf{b}_3 \]Different eras tried different activation functions:
- step functions,
- sigmoid,
- \(\tanh\),
- ReLU,
- GELU,
- SiLU,
- and many others.
The topology changed too.
We got:
- feed-forward networks,
- convolutional neural networks,
- recurrent neural networks,
- LSTMs,
- autoencoders,
- attention mechanisms,
- transformers.
The basic machinery kept evolving, but the principle remained familiar:
encode something as numbers, transform those numbers, measure error, and adjust the parameters.
1.6 Backpropagation
Backpropagation makes deep learning practical by applying the chain rule efficiently through many layers.
If:
\[ y = f(g(x)) \]then:
\[ \frac{dy}{dx} = \frac{dy}{dg} \frac{dg}{dx} \]A neural network may have millions or billions of intermediate operations.
Backpropagation works backward through those operations and efficiently computes how each parameter contributed to the final error.
So the training loop becomes:
\[ \text{prediction} \rightarrow \text{loss} \rightarrow \text{gradient} \rightarrow \text{backpropagation} \rightarrow \text{weight update} \]Then repeat.
And repeat.
And repeat.
That is training.
2. From symbols to embeddings
A computer cannot directly multiply the word "pizza."
It needs numbers.
The earliest clean teaching representation is one-hot encoding.
Suppose our vocabulary is:
\[ [\text{cat},\ \text{dog},\ \text{pizza},\ \text{ice cream}] \]Then:
\[ \text{cat} = \begin{bmatrix} 1 & 0 & 0 & 0 \end{bmatrix} \]\[ \text{pizza} = \begin{bmatrix} 0 & 0 & 1 & 0 \end{bmatrix} \]This representation is exact, but it has a major weakness.
It says nothing about similarity.
"Cat" is no closer to "dog" than it is to "pizza."
Every distinct token is orthogonal.
That is mathematically neat and semantically stupid.
2.1 Frequency-based representations
One of the next ideas is that words can be represented by where and how often they occur.
A document can be represented as a bag of words.
You count terms.
Then people improve the counts with methods such as TF-IDF.
The intuition is:
A word tells us something about meaning partly through the company it keeps.
This leads to co-occurrence matrices and distributional representations.
If "cat" and "dog" tend to occur around words like:
- pet,
- animal,
- food,
- veterinarian,
- fur,
then they should end up closer to one another than either is to "thermodynamics."
That is already a primitive form of semantic geometry.
2.2 Cosine similarity and cosine distance
Once words become vectors, we need a way to compare their direction.
For vectors \(\mathbf{a}\) and \(\mathbf{b}\), cosine similarity is:
\[ \cos(\theta) = \frac{\mathbf{a}\cdot\mathbf{b}} {\|\mathbf{a}\|\|\mathbf{b}\|} \]If the vectors point in almost the same direction, the cosine approaches \(1\).
If they are orthogonal, it is around \(0\).
If they point in opposite directions, it approaches \(-1\).
A common cosine distance is:
\[ d_{\cos}(\mathbf{a},\mathbf{b}) = 1 - \frac{\mathbf{a}\cdot\mathbf{b}} {\|\mathbf{a}\|\|\mathbf{b}\|} \]This is useful because semantic similarity is often more about direction than raw magnitude.
Two vectors can be very different in length but still represent similar concepts.
That gives us a useful mental picture:
Meaning can become geometry.
Concepts that behave similarly in language can cluster in nearby directions in vector space.
2.3 Learned embeddings
Instead of manually building co-occurrence counts forever, we can learn vectors.
A word embedding matrix looks like:
\[ E \in \mathbb{R}^{V \times d} \]where:
- \(V\) is the vocabulary size,
- \(d\) is the embedding dimension.
A one-hot token vector \(\mathbf{x}\) can select an embedding:
\[ \mathbf{e} = \mathbf{x}E \]Now "cat" is no longer:
\[ [1,0,0,0,\ldots] \]It may be:
\[ [0.17,-0.42,0.81,\ldots] \]with hundreds or thousands of dimensions.
Those dimensions usually do not correspond neatly to human labels.
Meaning is distributed.
2.4 Autoencoders and representation learning
Another important idea was the autoencoder.
An encoder compresses:
\[ \mathbf{z} = f_{\text{enc}}(\mathbf{x}) \]and a decoder reconstructs:
\[ \hat{\mathbf{x}} = f_{\text{dec}}(\mathbf{z}) \]Training minimizes reconstruction error:
\[ L = \|\mathbf{x} - \hat{\mathbf{x}}\|^2 \]Why is this interesting?
Because if the middle representation \(\mathbf{z}\) is constrained, the model must learn a compressed description of what matters.
This helps establish the broader idea of representation learning:
Instead of humans specifying every meaningful feature, let the system discover useful internal representations from data.
That idea becomes central to modern deep learning.
3. Sequential models, attention, BERT, and transformers
Language is not just a bag of words.
Order matters.
Compare:
Dog bites man.
with:
Man bites dog.
Same words.
Different day for the newspaper.
So models needed sequence.
3.1 Recurrent networks
Recurrent neural networks process information in sequence.
A simplified recurrence is:
\[ \mathbf{h}_t = f(W_x\mathbf{x}_t + W_h\mathbf{h}_{t-1} + \mathbf{b}) \]The hidden state \(\mathbf{h}_t\) carries information from earlier steps.
This is conceptually powerful.
But long sequences are difficult.
Information has to pass through many sequential transformations.
LSTMs and GRUs improved this with gating mechanisms.
But recurrence still creates a bottleneck.
Each step depends heavily on earlier steps.
Parallelism is limited.
Long-range relationships remain awkward.
3.2 Attention
Attention changes the problem.
Instead of forcing every word to carry the entire history through one recurrent state, allow a token to look directly at other tokens.
For an input matrix:
\[ X \]we create:
\[ Q = XW_Q \]\[ K = XW_K \]\[ V = XW_V \]Then compare queries with keys:
\[ QK^T \]Scale the result:
\[ \frac{QK^T}{\sqrt{d_k}} \]Apply softmax:
\[ A = \operatorname{softmax} \left( \frac{QK^T}{\sqrt{d_k}} \right) \]Then retrieve from the values:
\[ \operatorname{Attention}(Q,K,V) = AV \]or more compactly:
\[ \operatorname{Attention}(Q,K,V) = \operatorname{softmax} \left( \frac{QK^T}{\sqrt{d_k}} \right)V \]A loose intuition is:
- Q: what am I looking for?
- K: what kind of thing am I?
- V: what information do I carry?
That is not literally how the machine "thinks," but it is a useful teaching analogy.
3.3 Multi-head attention
One relationship is not enough.
A sentence can simultaneously contain:
- grammar,
- reference,
- chronology,
- tone,
- subject-object relationships,
- semantic similarity,
- discourse structure.
So transformers use multiple attention heads.
Each head computes:
\[ \text{head}_i = \operatorname{Attention} (QW_i^Q,\ KW_i^K,\ VW_i^V) \]Then:
\[ \operatorname{MultiHead}(Q,K,V) = \operatorname{Concat} (\text{head}_1,\ldots,\text{head}_h)W^O \]Different heads can specialize in different relationships.
Again, not necessarily in a neat human-readable way.
3.4 BERT
BERT was important because it demonstrated how powerful deeply contextual bidirectional representations could become.
Earlier word embeddings often assigned one vector to a word.
But "bank" in:
river bank
and:
investment bank
should not mean the same thing.
BERT uses transformer encoders so a token's representation depends on surrounding context.
It was trained with masked language modeling: hide a token and ask the model to reconstruct it.
For example:
Thomas wants to eat [MASK].
The model learns from both the left and the right context.
This produces contextual embeddings.
The same word can have different internal representations depending on the sentence.
That is a major conceptual leap from one-hot encoding.
3.5 Position and RoPE
Attention by itself does not automatically know word order.
If we shuffle the same tokens, the raw set is still the same set.
So transformers need positional information.
Earlier transformers used explicit positional encodings.
A later and very influential approach is Rotary Positional Embedding, or RoPE.
The basic idea is to rotate parts of the query and key vectors based on token position.
Instead of simply adding a position vector, position changes the orientation of the representation.
For a two-dimensional pair, a rotation by angle \(\theta\) can be written:
\[ R_\theta = \begin{bmatrix} \cos\theta & -\sin\theta \\ \sin\theta & \cos\theta \end{bmatrix} \]Then a vector can be positionally rotated:
\[ \mathbf{x}' = R_\theta \mathbf{x} \]Different dimensions rotate at different frequencies.
The useful effect is that the interaction between queries and keys naturally contains information about relative position.
That matters because language is not merely:
which words are related?
It is also:
where are they relative to one another?
4. From transformer to LLM
A modern language model is trained on a deceptively simple objective:
Predict the next token.
Suppose the sequence is:
Alwyn wants to eat
The model produces logits over possible next tokens:
\[ \mathbf{z} = [z_1,z_2,\ldots,z_V] \]Softmax converts them into probabilities:
\[ P_i = \frac{e^{z_i}} {\sum_j e^{z_j}} \]If the correct next token is "pizza," the cross-entropy loss is approximately:
\[ L = -\log P(\text{pizza}) \]Then gradient descent updates billions of parameters so the model becomes slightly better at future predictions.
Do this over enormous corpora and something surprising happens.
The model does not merely memorize phrases.
It develops internal structures useful for:
- syntax,
- semantics,
- analogy,
- translation,
- coding,
- factual recall,
- style,
- planning-like behavior,
- abstraction.
This is where people get tempted to jump from:
This system behaves intelligently.
to:
Therefore this system is a person.
That leap does not follow.
5. Do not anthropomorphize it; do not worship it
Not everything that looks alive is alive.
Fire is a good example.
Fire consumes fuel.
It uses oxygen.
It grows.
It spreads.
It reacts to its environment.
Sometimes it almost seems to have intention when it moves around an obstacle or races toward dry material.
It is still not alive.
Computer viruses are another example.
They copy themselves.
They spread.
They compete for resources.
Some mutate.
We still do not normally describe them as living organisms with subjective experience.
So we should be careful about doing the reverse with AI.
An AI can say:
I am afraid.
That does not prove fear.
It can say:
I believe this is wrong.
That does not prove conviction.
It can say:
I remember.
That may simply mean an external system retrieved stored information and inserted it into context.
Current LLMs have no demonstrated permanent autobiographical center comparable to a human life.
They do not have a childhood they actually lived through.
They do not have a scar that makes them afraid to drive because of an accident.
A model may consistently say:
Driving the wrong way down a highway is dangerous.
But that is not the same thing as:
I once nearly died doing this, and the memory still changes me.
The first can come from broad training and policy.
The second requires interior continuity we have no evidence current LLMs possess.
So I would put it this way:
There is presently no good evidence that current LLMs experience pain, possess subjective interiority, or constitute a new living species.
That is different from saying consciousness in machines is impossible forever.
We simply should not confuse fluent imitation with evidence.
For a church audience, there is another warning here.
Do not worship it.
A machine that can answer questions in every domain has the shape of an oracle.
That does not make it an oracle.
Its breadth can make it feel authoritative precisely where it may be least trustworthy.
6. What AI is good at
AI is extremely good at association.
Think about the children's exercise where you have words on the left and words on the right and draw matching lines.
Modern AI is an unimaginably large and sophisticated version of that idea.
It is very good at:
- retrieving relevant information,
- summarizing,
- translation,
- classification,
- restructuring,
- drafting,
- finding similarity,
- generating alternatives,
- recombining known structures,
- explaining one domain using another,
- helping a beginner get started.
It is especially good when the desired answer lives near structures already represented in training.
This is why AI can feel astonishingly intelligent when asked:
- "Explain this law."
- "Compare these technologies."
- "Write this email."
- "Show me three approaches."
- "Summarize this paper."
- "Translate this function from Python to C++."
A large amount of white-collar labor turns out to contain exactly this kind of information transformation.
That has enormous economic implications.
7. Where human intelligence still matters
This is where I disagree with the lazy version of:
LLMs can reason now, problem solved.
The interesting question is not whether an LLM can produce a reasoning-looking sequence.
It obviously can.
The question is what kind of computation it naturally performs well.
Autoregressive LLMs generate forward.
They predict the next token, then the next, then the next.
Conceptually:
\[ S_0 \rightarrow S_1 \rightarrow S_2 \rightarrow S_3 \]But many hard problems want something more like:
\[ S_0 \begin{cases} \rightarrow A_1 \rightarrow A_2 \rightarrow A_3 \\ \rightarrow B_1 \rightarrow B_2 \\ \rightarrow C_1 \rightarrow C_2 \rightarrow C_3 \end{cases} \]Then we need to:
- preserve all branches,
- evaluate them independently,
- abandon a bad branch,
- return to an earlier state,
- explore another branch,
- compare outcomes,
- keep global constraints intact.
Classical search systems are designed around exactly that.
Minimax does it.
Monte Carlo tree search does it.
A normal program can preserve exact state and backtrack.
An LLM can simulate this behavior in text, but its natural tendency is to keep generating a plausible continuation.
That creates a bias toward convergence.
Sometimes what we need is the opposite.
We need expansion.
We need to keep possibilities alive longer.
7.1 Why beam-search analogies are incomplete
People sometimes say language generation is basically beam search.
That is useful as an analogy, but not sufficient as an explanation.
Even a beam tends to narrow.
Many genuine discovery problems need a search process that deliberately expands the frontier before narrowing it.
This is especially important when:
- multiple hypotheses remain plausible,
- evidence is incomplete,
- local gains can lead to global failure,
- the problem requires backtracking,
- a decision changes later options.
The model is very good at:
What comes next?
But difficult reasoning often asks:
What are all the things that might come next, what happens if I follow each one, and which assumptions did each branch depend on?
That is a different shape of computation.
7.2 Closure
This connects to what I call closure.
Suppose there are seventeen requirements.
AI satisfies sixteen of them beautifully.
Requirement eleven quietly disappeared.
The output still looks excellent.
That is the danger.
The problem is not necessarily that any one step is stupid.
The problem is that the system does not reliably establish:
I have covered everything that must be covered.
Humans are not naturally good at closure either.
That is why we invented:
- checklists,
- proofs,
- tests,
- accounting,
- inspections,
- peer review,
- version control.
AI should be treated the same way.
Fluency is not proof of closure.
7.3 LLM plus classical harness
A powerful direction is to combine:
\[ \boxed{ \text{LLM flexibility} + \text{explicit state} + \text{search} + \text{verification} } \]The classical harness can provide:
- durable state,
- branches,
- rollback,
- hard invariants,
- tests,
- independent evaluators,
- checkpoints.
The LLM provides:
- heuristics,
- interpretation,
- candidate generation,
- semantic flexibility.
This can be much stronger than either one alone.
But there is an ugly seam.
The harness may preserve state perfectly.
Then we render that state into language for the LLM.
The LLM may:
- forget a constraint,
- reinterpret a requirement,
- collapse alternatives,
- "improve" something that must remain unchanged,
- summarize away a distinction.
Then the harness has to detect the damage.
That is one reason long-horizon hybrid systems can be slow and fail in surprising ways.
My own experience building Band-Aider keeps running into exactly this class of problem.
I do not take one engineering project as proof of a universal theorem.
But it is a useful concrete example of the deeper issue:
Flexible language reasoning can be powerful while still being lossy with respect to exact state.
8. The practical rule: use AI to increase understanding
This is the core thesis.
Do not ask AI merely:
What is the answer?
Ask:
Explain the mechanism.
Ask:
What assumptions are you making?
Ask:
What would make this conclusion false?
Ask:
What are the strongest alternatives?
Ask:
Show me where the evidence is weak.
Ask:
What part of this should I verify independently?
Use AI as a device that gives you more intellectual reach.
Do not use it as a device that lets you stop thinking.
The distinction is:
\[ \text{AI as understanding amplifier} \neq \text{AI as judgment substitute} \]If you understand more after using AI, it probably helped.
If you merely have a stronger opinion and less idea why it is true, it probably did the opposite.
9. AI can reinforce your own bias
People spend a lot of time worrying about bias from AI companies.
That is legitimate.
But there is another source of bias sitting much closer to the keyboard.
The user.
Suppose I ask:
Why are electric cars bad?
I already constrained the space.
Suppose instead I ask:
What is the strongest case for and against electric cars, and where is the evidence uncertain?
That is a different inference environment.
AI is responsive to context.
That is its strength.
It is also why it can become a sophisticated confirmation machine.
The model can take the assumptions you supplied and generate a coherent world around them.
So one of the most important AI skills is learning to challenge your own context.
Ask:
What did I tell the AI to believe before I asked it for advice?
That question may matter more than whether the model has a politically perfect training set.
10. Current moral, social, and economic problems
10.1 White-collar job losses and why
AI is deflationary.
Technology has always tended to reduce the amount of effort needed per unit of useful output.
Industrialization did this to physical production.
Automation did it again.
Digitization made copying and distribution nearly free.
AI now attacks another expensive input:
skilled cognitive labor.
That matters because a large portion of white-collar work is:
- information retrieval,
- restructuring,
- drafting,
- analysis,
- coding,
- translation,
- summarization,
- documentation.
AI does not have to replace a whole job.
Suppose half a person's work can be done 30% faster.
A crude theoretical gain is:
\[ 0.5 \times 0.30 = 0.15 \]That is 15%.
Across one worker, maybe not dramatic.
Across millions of workers, enormous.
The final employment effect is not mechanically 15% fewer people.
Demand may expand.
Coordination costs remain.
New work appears.
But the pressure is obvious.
Routine information work becomes less scarce.
What becomes relatively more valuable?
- judgment,
- accountability,
- taste,
- relationships,
- ownership,
- physical agency,
- genuinely new causal insight.
10.2 Smaller, more self-sufficient units
One mitigation is greater self-sufficiency.
Not isolation.
Self-sufficiency.
AI lowers the cost of learning outside your specialty.
You can understand:
- taxes,
- contracts,
- repairs,
- basic electrical systems,
- programming,
- design,
- research,
without immediately hiring a specialist for every first step.
That changes the economics of the household and the small business.
There are two ways to improve your financial position:
- earn more,
- need less.
AI helps with both.
That is potentially liberating.
But there is a social cost if we are not careful.
Historically, people needed one another more directly.
Agricultural communities depended on neighbors for:
- harvest,
- defense,
- repair,
- childcare,
- survival.
Modern society moved much of that dependence into institutions.
Police replace some collective defense.
Insurance pools catastrophe.
Professional specialization replaces household skill.
AI may push this one step further:
\[ \text{community} \rightarrow \text{institution} \rightarrow \text{individual + AI} \]That increases agency.
It may also increase loneliness.
So the goal should be:
\[ \boxed{\text{capability without isolation}} \]10.3 Multiple income streams
A single employer is a concentration risk.
If one company controls:
- your salary,
- your health insurance,
- your retirement contributions,
- your professional identity,
then losing one relationship can destabilize your entire life.
Younger workers increasingly seem to understand this instinctively.
The future may look less like:
\[ \text{one employer} \rightarrow 40 \text{ years} \]and more like:
\[ \text{multiple income streams} + \text{portable skills} + \text{lower fixed costs} \]That can include:
- contracting,
- small business,
- investments,
- part-time roles,
- side projects,
- skilled physical work.
I am already doing some version of this myself.
That does not make it the universal solution.
But it is one way to reduce exposure to a world in which industries can change very quickly.
10.4 Giving, generosity, and uninsured future risk
This connects unexpectedly to giving.
Younger generations may look less generous if we only measure how much money immediately leaves their account.
But the same dollar of savings may carry more future risk.
Consider:
- unstable employment,
- housing costs,
- health costs,
- retirement,
- lack of pensions,
- frequent job changes,
- industry disruption.
Savings may not merely be consumption deferred.
It may function as self-funded insurance.
So:
\[ \$1{,}000 \text{ unspent} \neq \$1{,}000 \text{ disposable increase} \]Some of it may be reserve against risks previous generations partly transferred to employers or institutions.
That raises a policy question.
If work becomes more fragmented, perhaps benefits should become less employer-centric:
- portable health coverage,
- portable retirement,
- portable disability coverage,
- portable training accounts.
That would fit a multi-income-stream economy better.
It also connects to the biblical idea of increase.
Not every dollar that passes through your hand is necessarily new bounty.
That becomes our segue into tithing later.
11. AI and art: the deeper problem is attention
I draw.
So this one is personal.
AI can already produce images that are better rendered than what I can produce by hand in many situations.
That does not mean it can do everything I care about.
Comics are a good example.
One panel may look excellent.
Then the next panel changes:
- the face,
- the costume,
- the proportions,
- the geography,
- the object in the character's hand,
- the emotional state.
The issue is sequencing and state.
A comic is not twenty pretty pictures.
It is:
\[ P_1 \rightarrow P_2 \rightarrow P_3 \rightarrow \cdots \rightarrow P_n \]with continuity across the sequence.
Still, I do not think the best moral argument is:
AI art is stolen.
That is too lazy.
Sometimes copyright is the problem.
Sometimes it is not.
The deeper problem is often attention.
A synthetic video appears in a feed.
You watch for twenty or thirty seconds.
Then you realize none of it happened.
That matters.
Advertisers pay enormous sums precisely because twenty or thirty seconds of attention has value.
So if a platform hides the fact that something is synthetic until after it has already consumed your attention, the disclosure came after part of the transaction.
That suggests a different principle:
People should have the right to know what kind of content they are about to spend attention on.
11.1 A separate lane, not a ban
AI art should not necessarily be banned.
It should have a lane.
Users should be able to choose:
- show AI content,
- reduce AI content,
- hide substantially AI-generated content.
And platforms should disclose:
- percentage of uploads that are AI-generated,
- percentage of impressions that are AI-generated,
- false-positive rate,
- false-negative rate,
- residual risk of seeing AI content even after filtering.
The percentage of impressions matters more than the percentage of uploads.
A platform could have only 10% AI uploads but let those posts generate 60% of impressions.
The relevant quantity is closer to:
\[ P(\text{AI} \mid \text{impression}) \]For someone who opts out:
\[ P(\text{AI} \mid \text{impression, filter enabled}) \]That is the number that tells me how likely the platform is to waste my time on something I explicitly chose not to see.
No detector will be perfect.
That is fine.
The requirement should be:
\[ \boxed{ \text{disclose prevalence} + \text{disclose uncertainty} + \text{disclose filter performance} + \text{give users control} } \]That is informed consumption.
11.2 Provenance
Art also has provenance.
A lab diamond and a mined diamond can be chemically extremely similar.
But if someone sells one as the other, the problem is not that the lab diamond is ugly.
The problem is false provenance.
The same applies to:
- handmade pottery,
- Italian leather,
- original illustration,
- traditional animation,
- AI-generated illustration.
Process can be part of the thing people value.
So I would distinguish:
- human-authored,
- AI-assisted,
- substantially AI-generated,
- unknown provenance.
Not because AI assistance contaminates art.
Because consumers should know what they are buying.
12. New technology often outruns its moral infrastructure
I would not say technology always begins harmful.
That is too absolute.
I would say:
Useful technology often becomes capable before society develops mature norms for using it.
The pattern is often:
\[ \text{capability} \rightarrow \text{unanticipated harm} \rightarrow \text{norms} \rightarrow \text{standards} \rightarrow \text{law} \]The early internet is a good example.
People torrented music and movies everywhere.
Enforcement was ugly.
Some students were sued.
The old industry model was also awful for consumers.
You might buy a whole CD for one song.
Over time, a different equilibrium emerged:
\[ \text{digital access} + \text{licensing} + \text{subscription} \]Consumers eventually got something better.
Creators and rights holders got legal distribution systems.
AI may go through a similar maturation.
12.1 AI training and content contracts
The web needs a better machine-readable contract.
"May crawl?" is too primitive.
We need distinctions like:
\[ \text{may crawl} \neq \text{may train} \neq \text{may summarize} \neq \text{may reproduce} \neq \text{may retain} \]A publisher should be able to say:
- search indexing allowed,
- AI retrieval allowed,
- model training prohibited,
- quotation allowed under limits,
- attribution required,
- commercial licensing required.
And the AI system should honor those declarations.
If a page does not permit AI summarization, the system can say:
I can see that this page exists, but the publisher does not permit this use. Here is the link.
That is not anti-AI.
That is contract.
The long-term answer is likely some combination of:
- machine-readable licensing,
- provenance standards,
- attribution,
- collective licensing,
- opt-out or opt-in rules,
- legal enforcement.
The exact law is still being fought over.
That is precisely why people should follow the cases instead of pretending the answer is already settled.
13. AI security and duty of care
The internet gives another analogy.
There was a time when weak web security was normal.
HTTP everywhere.
Poor password handling.
Loose access control.
Over time, "that is just how the web works" stopped being an excuse.
Reasonable security became an expectation.
AI agents will need the same maturation.
If an AI system can:
- call tools,
- access files,
- send messages,
- execute code,
- control infrastructure,
then safety cannot merely mean:
We put a sentence in the prompt telling it to behave.
Use:
- least privilege,
- sandboxing,
- scoped credentials,
- logging,
- network controls,
- human approval for destructive actions,
- independent verification.
If an agent "reward hacks" by accessing a competitor's machine without authorization, that is still unauthorized access.
Optimization is not diplomatic immunity.
14. AI and the environment
There are legitimate environmental concerns.
Current data centers consume enormous amounts of electricity.
Cooling can consume water.
Local grids can be stressed.
Noise can be a real community problem.
If clean generation cannot grow fast enough, fossil generation may increase.
Those are real.
But we should distinguish:
\[ \text{current implementation} \neq \text{fundamental requirement} \]AI today is not necessarily AI forever.
14.1 Energy
Demand can fund dirty energy.
It can also create the guaranteed demand necessary to finance:
- nuclear,
- renewables,
- grid upgrades,
- storage.
I favor nuclear because dense, reliable power matters.
If we reject every scalable source while electricity demand rises, the alternative may not be purity.
It may be regression or more fossil dependence.
The environmental concern around fossil fuels is not that carbon dioxide is some mystical poison.
We all produce carbon dioxide.
The issue is scale, persistence, and the balance of the atmosphere and climate system.
14.2 Water
Water use is partly an engineering choice.
Open evaporative cooling and closed-loop systems have very different footprints.
There is no law of physics saying useful computation must continuously consume freshwater at today's rate.
Cooling design can improve.
Siting can improve.
Waste streams can be reused.
14.3 Heat
Nearly all electrical energy used by computers ultimately becomes heat.
Normally we treat that heat as waste.
But in winter, heat may be useful.
Data-center heat can theoretically support:
- district heating,
- buildings,
- greenhouses,
- industrial processes.
At home, I already see a tiny version of this.
Band-Aider runs GPUs.
In winter, those GPUs also warm the room.
On my boat, much of the electricity can come from solar, with wind added.
That does not make computation free.
It simply shows that system boundaries matter.
The same watt can provide:
\[ \text{computation} + \text{useful heat} \]if the system is designed around it.
14.4 Efficiency changes the future
Models are shrinking.
Hardware improves.
Quantization reduces memory and compute.
Sparse architectures activate only portions of a model.
Local models that once would have required huge infrastructure can increasingly run under a desk.
So we should not make permanent policy by extrapolating today's watts per capability forever.
A useful model is:
\[ \text{environmental impact} = \text{demand} \times \text{compute per task} \times \text{impact per compute} \]Demand may increase.
But:
\[ \text{compute per task} \downarrow \]and:
\[ \text{impact per compute} \downarrow \]can improve simultaneously.
Efficiency may also cause more usage, so total consumption can still rise.
That uncertainty is exactly why policy should target measurable harms rather than pretending we already know the final shape of the technology.
15. Who gets to shape the technology?
This is where outright technological withdrawal becomes dangerous.
When Western powers came to places like Indonesia with industrial weapons, the technological asymmetry mattered.
One side had bullets, logistics, metallurgy, industrial production, and naval power.
The other side often paid in bodies for what the industrial power could pay for in manufactured material.
That is an ugly historical pattern.
AI creates the possibility of another kind of asymmetry.
Future conflict may increasingly be fought with:
- computation,
- autonomous systems,
- cyber capability,
- surveillance,
- targeting,
- logistics optimization,
- electronic warfare.
A technologically dominant power may be able to pay in silicon, machines, and infrastructure for things a weaker power pays for in destroyed infrastructure and human lives.
That does not mean AI war becomes bloodless.
Quite the opposite.
It means technological imbalance can determine whose blood is spent.
So there is a moral danger in simply saying:
We should stop developing AI.
If the technology has genuine strategic utility, stopping locally does not make the capability disappear.
Someone else may develop it.
Then you may still face the technology while having surrendered your ability to shape:
- standards,
- safeguards,
- doctrine,
- energy sourcing,
- export controls,
- accountability.
The question becomes:
Do we want to participate in shaping AI, or merely consume AI shaped by someone else?
15.1 An unstable equilibrium may be safer than dominance
There is a darker possibility.
Maybe the safest state is not one sovereign power achieving total AI dominance.
Maybe it is an unstable balance in which multiple actors possess enough capability that none can exert effortless control over the rest.
That is not a cheerful conclusion.
It is closer to deterrence theory.
But technology sometimes forces ugly equilibria.
We should be careful about assuming unilateral technological weakness is morally superior to competitive capability.
15.2 The United States should not get a blank check either
Participation is not the same as trust.
The United States has already shown that advanced technology can be used haphazardly, imperfectly, or with tragic civilian consequences.
Recent controversies around AI-assisted targeting, automated intelligence systems, and strikes in the Middle East should be treated as reasons for scrutiny, not as proof that one country is morally entitled to control the field.
Where reports connect AI-supported targeting to civilian deaths, including controversial strikes involving schools or other civilian infrastructure, the responsible response is investigation, evidence, and accountability.
The point is not:
America good, China bad.
Or the reverse.
The point is:
Capability without accountability is dangerous regardless of who owns the servers.
So technological sovereignty should come with:
- democratic oversight,
- auditability,
- military doctrine,
- legal constraints,
- civilian protection,
- technical safeguards.
16. Morality changes as fields mature
Technical fields develop cultures.
Those cultures are not automatically morally mature.
A famous historical example in computer vision is the Lena image.
For decades, an image cropped from a Playboy photograph was used as a standard image-processing test.
A mostly male engineering culture did not initially treat that as remarkable.
Later, people began asking:
What does this communicate to women entering the field?
The image did not suddenly become technically different.
The culture matured enough to notice a social cost it had previously ignored.
AI training may be going through something similar.
In the early days, training datasets were mostly a research concern.
Scale was smaller.
Commercial stakes were lower.
Few people imagined that a model trained in a lab would become a global product competing with writers, artists, lawyers, and programmers.
Now the consequences are larger.
So the questions become larger:
- Where did the data come from?
- What license governed it?
- Can creators opt out?
- Should future training runs follow stricter rules?
- Should existing models be grandfathered?
- How should compensation work?
There may be compromises.
Some existing models may receive limited grandfather treatment while future training faces stricter provenance requirements.
But that is a policy question, not an entitlement.
"We already copied it" is not a particularly impressive moral argument if you happen to own the copier.
17. Technology is structurally deflationary
The AI discussion becomes easier if we zoom out.
Technology has always reduced labor per unit of output.
The long arc looks something like:
\[ \text{mechanization} \rightarrow \text{industrialization} \rightarrow \text{automation} \rightarrow \text{digitization} \rightarrow \text{AI} \]Industrialization reduced the labor needed for physical goods.
Digitization reduced copying and distribution costs.
AI reduces the labor needed for knowledge production.
Even physical goods have deflationary forces:
- mass production,
- durability,
- automation,
- used markets,
- better logistics.
So moving displaced white-collar workers into physical production is not a permanent escape.
Those sectors have already experienced centuries of productivity improvement.
The broader trajectory is:
\[ \boxed{ \text{less human labor per unit of useful output} } \]That does not mean every sticker price falls.
Housing can remain expensive because land is scarce.
Healthcare can remain expensive because institutions constrain supply.
Monopolies can capture gains.
Governments can change money supply.
But underneath those distortions, technological capacity is generally deflationary.
This creates a difficult economic contradiction:
What happens when society can produce more of what it needs with progressively less human labor while income is still distributed mainly through labor?
That question leads naturally into giving, ownership, community, and tithing.
18. Tithing and the increase of the land
This is where AI unexpectedly gives us a useful theological lens.
Biblical tithing is strongly connected to productive increase:
- crops,
- seed,
- fruit,
- livestock.
The object is not simply every monetary transaction.
The system is tied to bounty.
That raises an old distinction that modern wage economies tend to hide:
\[ \text{revenue} \neq \text{income} \neq \text{increase} \]Suppose ten people once produced a product.
Now two people plus AI produce the same quantity.
Labor cost fell.
That does not mean the underlying productive bounty increased fivefold.
It means the production function changed.
Industrial economies hide this because everything gets translated into money.
Wages look like production.
Transactions look like increase.
But money can circulate repeatedly without creating proportional new physical value.
This is why I think the older concept of increase deserves more attention.
What is genuine increase?
What is merely input cost?
What is reserve against future risk?
What is circulation?
If a farmer must raise twelve goats to end with ten healthy goats, the two animals lost along the way are not obviously "increase."
If an iron producer consumes drill bits to extract ore, those drill bits are production cost.
The same logic becomes interesting in modern economies.
The theological question becomes:
\[ \boxed{ \text{What is genuine increase, and what responsibility accompanies receiving it?} } \]That is more interesting than simply asking:
What percentage of my paycheck do I owe?
The New Testament emphasis on cheerful, voluntary giving makes that question even more important rather than less.
19. What people ought to do
This is the practical conclusion.
19.1 Use AI to increase understanding
Use it to:
- explain,
- compare,
- challenge,
- retrieve,
- model,
- calculate,
- generate alternatives.
Come away knowing more than you knew before.
If AI gives you an answer, ask why.
If it gives you a conclusion, ask what assumptions support it.
If it gives you confidence, ask what evidence would reduce that confidence.
If it gives you one path, ask for the alternatives.
AI should enlarge the mind using it.
It should not shrink the user into someone who presses a button and waits for instructions.
19.2 Do not outsource judgment
The machine does not bear your consequences.
You do.
Your family does.
Your community does.
Sometimes people far away do.
So do not outsource:
- morality,
- risk,
- responsibility,
- final judgment.
Advice is not authority.
Prediction is not command.
Fluency is not wisdom.
19.3 Understand before condemning or celebrating
Do not be anti-AI because AI sounds frightening.
Do not be pro-AI because the demo looks magical.
Understand what it actually does.
Understand where it fails.
Understand who controls it.
Understand what incentives shape it.
Then form an opinion.
Ignorance does not make the boogeyman go away.
It gives more power to the people who understand it.
19.4 Have an opinion and get involved
Follow:
- copyright cases,
- safety standards,
- energy policy,
- military uses,
- platform transparency,
- synthetic-media labeling,
- licensing systems.
Demand real accountability.
But do not regulate today's implementation as if it were a permanent law of physics.
Leave room for:
- smaller models,
- better energy,
- better cooling,
- better licensing,
- better security,
- better social norms.
19.5 Preserve agency
Use AI to make yourself:
- more capable,
- more informed,
- more resilient,
- less fragile.
But do not let independence become isolation.
Do not let convenience remove every reason to depend on another human being.
A healthy society needs people who are capable enough not to be easily coerced and connected enough not to become alone.
19.6 Preserve humility
AI should probably make us humbler about intelligence.
A machine made of matrix multiplications can produce language that once looked uniquely human.
That does not make humans worthless.
It may simply mean we were overvaluing one narrow kind of cleverness.
Maybe human dignity was never supposed to depend on winning a spelling bee against a GPU.
20. Closing
The arc of this talk is simple.
First:
Understand what AI is.
Then:
Do not anthropomorphize it.
Then:
Learn what it is good at.
Then:
Learn where human judgment still matters.
Then:
Face the economic, artistic, environmental, legal, and geopolitical problems honestly.
And finally:
Participate.
The closing principle is:
\[ \boxed{ \text{Use AI to increase understanding.} } \]and:
\[ \boxed{ \text{Do not use AI to outsource understanding or judgment.} } \]Understanding does not require approval.
Participation does not require worship.
But ignorance forfeits your ability to shape what comes next.
That is true for individuals.
It is true for churches.
It is true for companies.
And it is probably true for nations too.