Marvin Minsky   1994

본 문서에서는 추론을 하는 전문가의 전문지식이 필연적으로 오류 (fallacy)를 가지며 우리의 조상들이 생존을 위해 돌발사고 로부터 피해가는 방법을 배워왔듯이 (acquired learning machine), 복잡한 추론에 따른 오류를 피해가는 방법들을 터득하여 가야한다는 것을 강조한다.

"전문가란 생각할 필요도 없이, 아는 사람을 의미한다 "--Frank Lloyd Wright

Abstract: We tend to think of knowledge in positive terms -- and of experts as people who know what to do. But a 'negative' way to seem competent is, simply, never to make mistakes. How much of what we learn to do -- and learn to think -- is of this other variety? It is hard to tell, experimentally, because knowledge about what not to do never appears in behavior. And it is also difficult to assess, psychologically, because many of the judgments that we traditionally regard as positive -- such as beauty, humor, pleasure, and decisiveness -- may actually reflect the workings of unconscious double negatives.

이러한 경향은 AI 의 연구로 나타난 rule-based expert system에서 표현된다. 사실상 그들의 지식의 전부가 " IF X 가 발생하면, Y 가 수행된다."라고하는 positive rule 들로 코드되었다. 그러나 이것은 많은 전문지식을 상실한다. 확실히 사람들이 해야할 것을 알아야 하는 것도 필요하지만 그러나 또한 해서는 안되는 것을 알 것도 요구한다. " IF 당신이 절벽에 가까이 있다면, THEN 그쪽을 향해 걸어가지 마라"  전문가는GOAL을 얻는 방법을 알 뿐만 아니라 재앙을 피하는 방법도 알아야 한다. 때때로 우리는 돌발사고에 대해 정면대응하는 positive 태도를 취할수도 있지만 대개는 문제를 일으키는 원인이 되는 행동을 피함으로써 해결한다. 본 문서에서는 많은 인간의 지식이 negative 하다는 것을 주장한다.

This inclination is expressed in the "rule-based expert systems" that emerged from research in AI. Virtually all of their knowledge is encoded as positive rules: "IF X happens, DO Y." But this misses much of expertise. Certainly, competence often requires one to know what one must do -- but it also requires you to know what not to do. "IF you are close to a precipice, DON'T walk toward it." An expert must know both how to achieve goals and how to avoid disasters. Sometimes we can take positive measures against accident -- but mostly we do it by avoiding actions that might cause trouble. This essay argues that much of human knowledge is negative.

 똑 같은 것이 사고 (thinking) 에도 적용된다. 효율적으로 사고하기 위해서 우리는 사고하지 않는 많은 것에 대하여 알아야만 한다. 그렇지 않으면 우리는 bad idea를 얻거나 또는 idea를 얻는데 너무 오래걸린다. 이것은 많은 논리적 과제들을 야기한다.

And the same applies to thinking, as well. In order to think effectively, we must "know" a good deal about what not to think! Otherwise we get bad ideas -- and also, take too long. This raises a number of theoretical issues:

왜 negative knowledge 가 중요한가?  세상은 살아가기에 위험한 곳이다. 예를들면 생물학자들은 대부분의 돌연변이들은 해롭다고 얘기한다. 왜냐하면 각 동물들은 이미 돌연변이체의 공간에서 일종의 local optimum (local environment 와 관련하여) 가까이에 있기 때문이다. 어떤 정상에 가까이 있으면 대부분이 아래로 내려오기 마련이다.

Why is negative knowledge important?

The world is a dangerous place for life. For example, biologists tell us that most mutations are deleterious. This because each animal is already near a sort of local optimum (with regard to its local environment) in the space of mutational variants. And near the top of any hill, most steps go down.

그러나 왜 각 동물들은 local peak 에 있는 것일까? 단순히 진화 그자체는 정상에 오르도록 만들어진 일종의 learning machine 이기 때문이다. 모든 기존의 동물들은 후손들이 가지게 될 돌발사고들을 충분히 피할수 있게 만든 조상들이 있으며, 그 조상들은 독, 질병, 약탈자, 경쟁자, 기타 위험한 상황들을 피하도록 학습되도록 하는 기술들을 획득한 존재이다 (acquired machinery). 물론 우리는 positive goal을 배우고 그것을 성취하는 방법을 배우도록 진화했다. 그러나 여전히 우리 세계는 기회보다는 더 많은 위험을 가지고 있으며 우리의 최상의 목표는 죽지 않고 생존하는 것임에 틀림없다!  위험을 피하기 위해서는 많은 방법이 있으며 당신의 적으로부터 탈출하기위해서는 그들을 파괴하고, 조절하고, 피해야만 한다. 아마도 우리의 사회와 문화와 정부는 대부분의 흔한 돌발사고의 원인에 대항하여 보호해 주기 위한 negative goal 에서부터 유래된 것일것이다.

But why is each animal close to a local peak? Simply because evolution itself is a learning machine that is engineered to climb hills. All existing animals had ancestors that avoided enough accidents to have descendants, and those ancestors were just the ones that acquired machinery that enabled them to learn to avoid poisons, diseases, predators, competitors, and other dangerous situations. Of course we also evolved to learn positive goals and ways to achieve them; still, to the extent that our world offers more perils than opportunities, our topmost goal must be -- don't get killed! There are many ways to avoid dangers You can escape your enemies by destroying, controlling , or evading them. Perhaps our societies, cultures, and governments themselves originated in negative goals, namely, for protection against the most common causes of accidents.

지능의 진화는 많은 새로운 기회를 가져다 주었지만 또한 우리에게 실패할 가능성도 커지게 하였다. 우리가 추론의 능력을 가지게 되자마자 곧 오류의 가능성에 직면하게 만들었다. 우리의 계획의 범위를 확대시킴에 따라 우리는 더 많은 복잡한 종류의 잘못에 빠지는 경향이 있다. 말하는 기술이 진화함에 따라 다른 사람의 나쁜 생각에 감염될 위험을 증가시키었다. 육체뿐만아니라 정신세계도 좋은 것보다는 더 나쁜 것을 포함할수 있다. 물론 대화는 다른 사람에게 좋고 나쁜 것에 대한 면역을 주는 생각들을 전파시킬수 있다.

The evolution of intelligence brought great new opportunities -- but also gave us great new ways to fail. As soon as we were capable of reasoning, we became susceptible to fallacies. As we extended the range of our plans, we fell prone to more intricate kinds of mistakes. As the arts of speech evolved, this increased the risk of infection by more bad ideas from other minds. The mental, as well as the physical world may also contain more bad than good. Of course, communication can also transmit ideas that give immunities to other, good and bad, ideas.

뇌는 많은 특별한 기능 (agencies)를 가지며 많은 다른 표현방식을 사용한다. 어떤 agencies는 연속적인 개념을 표현하기 위해 필기문서같은 구조를 사용하고, 다른 것은 트리구조의 자료구조를 가지고, 계층구조나 더 복잡한 구조를 위해 semantic network을 사용하고, 절차의 효율적인 수행을 위해서는 rule 의 production collection을 사용하고, 인과율에 대한 추론을 위해서는 trans-frame 구조를 사용한다. 신경에 대한 책의 목차만 대강 훓어보아도 뇌는 많은 해부학적인 영역과 그것을 연결하는 신경의 다발들을 포함하고 있는 것이 보일 것이다. 왜 뇌는 사고하기 위해서 그렇게 많은 방법들을 가져야 하는 것일까. 내가 생각하기에 그것은 우리가 세계에서 직면하는 많은 종류의 문제들을 위해서 작동하기 때문일 것이다.  각각의 problem solving 전략, 각각의 사고 방식, 각각의 지식표현들이 어떤 영역에서는 작동하지만 다른 영역에서는 실패한다. 연속해서, 우리가 축적하는 각각의 지식에 대해서, 우리는 그러한 지식베이스를 사용할 때와 사용하지 않을 때를 구분하는 지식들을 또한 필요로 한다.

The brain has many specialized agencies, In the later chapters of [SOM] I argue that these must use a variety of different representations. Some agencies might use script-like structures for representing sequential concepts and story-like exemplars. Others may use tree-like data-structures and/or semantic networks for hierarchical classifications and more complex structures; topographical arrangements for representing spatial and haptic situations, production-like collections of rules for efficient execution of procedures, and "trans-frame" like structures for reasoning about causality. A cursory glance at the index to a neurology book shows that the brain includes hundreds of anatomically distinct "regions" and bundles of fibers that interconnect them. Why should a brain have so many ways to do things. I think, because so single scheme will work for all the many kinds of problems that the world confronts us with. Each problem-solving strategy, each style of thinking, each knowledge-representation scheme -- each works in certain areas, but fails in other domains. Consequently, for each body of knowledge we accumulate, we also need knowledge about when to use that knowledge-base and when to not.

Some of my colleagues have argued that the brain is not a suitable basis for such a discussion, because it was never really designed to think. Surely (they say) most of that complexity could be avoided when we design such machines from scratch. Surely (some of them maintain) we can construct a single uniform, consistent, and effective logical systems to perform all kinds of commonsense reasoning. I doubt [1991 Minsky] this will be feasible, because consistency and effectiveness may well be incompatible. Other colleagues maintain that we should be able to construct large, uniform neural networks that can learn to do all that minds might need. I do not see much hope of this, because of fear that any very large such network would be prone to accumulate too many interconnections and become paralyzed by oscillations or instabilities. How could we stabilize such systems ? My answer is that one might have to provide a variety of alternative sub-systems, decoupled enough that if each part should fail from time to time, the rest could continue to function so that not all the system will all fail at once. This means that those parts must be suitably insulated from one another. There has been so little recognition of this problem in modern AI that perhaps we need a new term for it. Perhaps we need to breed researchers who can call themselves "Insulationists".

Of course, some insulationist functions already are encompassed by traditional learning theories. Because the most popular forms of neural networks and fuzzy logic can reduce as well as increase their weights, this could tend locally to eliminate 'detrimental' connections. But it is my feeling that although that sort of thing is formally possible, it is heuristically impractical. Instead, I maintain, effective systems will need to be provided with suitable architectures from the start. Perhaps we'll have to design each agency with appropriately engineered machinery to prevent our machines from getting stuck. For example, in [1985 Minsky, Ch.10] we proposed that a typical agency might be built to incorporate Seymour Papert's "Exclusion Principle,", so that when an agency develops a serious internal conflict, the subagents involved should be inhibited so that others can take over.

How much of human knowledge is negative?

We spend our lives at learning things, yet always find exceptions and mistakes. Certainty seems always out of reach. Except in worlds we invent for ourselves (such as formal systems of logic and mathematics) we can never be sure our assumptions are right, and must expect eventually to make mistakes and entertain inconsistencies. To keep from being paralyzed, we have to take some risks. But we can reduce the chances of accidents by accumulating two complementary types of knowledge:

We search for 'islands of consistency' within which commonsense reasoning seems safe .

We also work to find and mark the unsafe boundaries of those islands.

Both as cultures and as individuals, we learn to avoid patterns of thought reputed to yield poor results. In civilized communities, appointed guardians post signs to warn about sharp turns, thin ice, and animals that bite. And so do our philosophers, when they report to us their paradox-discoveries - those tales of Liars who admit to lying, and Barbers who shave all who do not shave themselves. These precious lessons teach us about which thoughts we shouldn't think; they are the intellectual counterparts to Freud's emotion-censors. It is interesting how frequently we find logically paradoxically nonsense to be funny, and when we come to jokes, we'll see why this such a humorous character. For when we look closely, we find that most jokes are concerned with taboos, injuries, and other ways of coming to harm - and logical absurdities can also potentially lead to harm.

I think we have neglected that second aspect -- of asking how experts manage to discern and defend the margins their islands of consistency. It is so hard to study what minds do /i[not] think that this may have placed that subject beyond the bounds of behaviorist psychology, because of its non-behavioral character. And introspective methods also fail because (like most of learning and reasoning) such processes are hidden from consciousness. Neil Agnew pointed out to me that this poses a problem for knowledge engineers -- those who would encode an expert's expertise. Presumably, experts have more effective censors than the rest of us -- but we can't rely upon their introspection to detect the work of their inhibitory agencies. Worse, perhaps as Freud proposed, our censors actively resist their exposure. Still, sometimes outsiders can see what in ourselves what we cannot, by noticing such nuances of behavior as avoiding, forgetting, displaying of temper, rationalizing, or citing only positive instances, etc. [See also Agnew & Brown (1986) for a discussion of confidence in reasoning.]

How can we implement negative knowledge?

One way is to divide the mind into parts that can monitor one another. For example, imagine a brain that consists of two parts, A and B. Connect the A- brain's inputs and outputs to the real world - so it can sense what happens there. But don't connect the B-brain to the outer world at all; instead, connect it so that the A-brain /i(is) the B-brain's world! Then A can see and act upon what happens in the outside world. On the other hand, B can only "see" and influence what happens inside A. This could be enough to help block some kinds of bad patterns of thinking in A.

If A is not making progress toward its goal, force it to review that goal.

If A seems to be repeating itself, make it stop and try something else.

If A does something B considers good, reinforce A's learning system.

If A is occupied with too much detail, then make it take a higher level view.

If A is not being specific enough, then make it focus on more details.

If A appears to be making things worse, suppress it in favor of another agency.

If A asks more than three 'whys' in a row, shift to another agency.

This sort of thing could be a step toward a more "reflective" mind-society. A B- brain could experiment with its A- brain, just as the A-brain can experiment with the real-world objects and the people that surround it. And just as A can try to predict and control what happens outside, B can try to predict and control what A will do. And even though B may have no concept of what A's activities mean in relation to the outer world, it is still possible for B to be useful in the sort of way that a counselor or management consultant can assess a client's mental strategy without having to understand all the precise details of that client's profession.

Emotions and NegExpertise

Negative knowledge is involves in many of the forms of thinking that we term 'emotional', notably those involved with humor, shame, fearful, and aesthetic appreciation. This machinery includes a variety of suppressors, critics, and inhibitors, some of which can inhibit not merely actions but entire strategies of thought. Thus, once one begins to look for it, one finds examples of negative knowledge in many activities that we usually see as positive. In the earliest theories about AI, for example, we emphasized the importance of heuristics for generating efficient search trees. This can be done either by pruning initially larger trees or by suppressing those branches right form the start -- that is, by not thinking of them in the first place. When you decide to leave a room, you don't even think of jumping out the window. Thus, a positive system forces us to generate and test, whereas a negative-based system could more efficiently shape the search space from the start. To do this efficiently, we would have to invent ways to compile each new search generator, perhaps on the basis of previously learned negative prototypes. To wait for inhibition during run time would consume more time.

This relates to what is commonly called creativity. It annoys me how frequently people suggest that the 'secret' of making creative machines might lie in providing some sort of random or chaotic kind of search generator. Nonsense! Certainly, there must be a source of variation -- but that can be supplied by all sorts of algorithmic generators. What distinguishes the performance of a 'smart' or 'creative' artist or problem-solver is not how many trials precede a success, but how few. So the secret lies not in disorderly search, but in pre-shaping the search space so as reduce the numbers of useless attempts. Of course, that's not the whole story. In order to establish an individuality, a creative modern artist must also generate some unconventional alternatives. Doing that may also involve un- suppressing some conventional censors. In any case, from the negative knowledge point of view, we might argue that often beauty is neither in the eye, nor even in the mind of the observer, but precisely the opposite: it may lie in the power to inactivate many of that observer's internal critics.

To explain this, let's consider the role of emotions in thought. It seems generally agreed upon that, on the whole, the positive emotions involve learning what to do, while the negative ones involve learning what not to do. But if so, then I suspect that many emotions that we normally see as 'positive' are actually not. For example, it seems to me that much of our celebrated sense of Beauty may be negative, no matter that we see it as positive. For when possessed by that emotion, many people seem to me to have suspended much of their normal question-asking machinery. When a person says, "how perfectly beautiful this is," they seem also to be saying, 'it is time to stop evaluating, selecting, and criticizing.' They often regard as hostile, requests to be asked to explain why they are attracted to it..

Humor is also usually seen as positive, no matter that the force of a joke is to say, "Don't even think about doing X," or "Don't take it seriously!" Most people are quite unaware that jokes are usually about things that one should not do, because they are prohibited, disgusting, or simply stupid.

Similarly, we tend to think of decision-making as positive. Yet the act of decision, which we often describe as an "act" of free will, is more of a NegAct by nature, because what seems consciously to be the moment of 'making' the decision is actually the moment of terminating the process of considering alternatives.

Perhaps it is the feeling of Pleasure that we consider most positive of all. Yet once we start to see the mind as not one, but a society of processes, then the most extreme pleasure can be seen instead as most negative. For it may mean merely that a certain process has seized control, and has managed to turn off most of the rest. Naturally, that makes it hard to think about anything else. Surely the most extreme form of the control of mental agencies can be seen in what we call mystical experience. For when this happens to a mind, it is like saying to oneself, "Now my problems are all solved. I know the Truth, and know that there is no need to question it, or seek confirming evidence. Stop thinking now, and let all Critics cease. "

We normally think of beauty, humor, pleasure, and decisiveness as positive; is it then paradoxical to claim the opposite? No, not at all -- because we're dealing with things complex enough to constitute 'double negatives'. Putting something in a folder labeled 'negative' can't keep it there, because we then can re-enclose it in a second sign-changing shell! Thus pleasure can seem positive to the agency now in control -- no matter that your other agencies are suffering under its yoke. Thus, enjoying something very much can mean that you've engaged machinery that (i) makes you think even more about that something and (ii) keeps you from thinking of other things.

Conclusion

We tend to think of knowledge in positive terms -- and of experts as people who know what to do. But a 'negative' way to seem competent is, simply, never to make mistakes. How much of what we learn to do -- and learn to think -- is of this other variety? How much of human competence is knowing methods for solving problems, and how much of it is knowing how to intercept and interdict unproductive lines of thought? It is hard to assess the importance of these, experimentally, because knowledge about what not to do never appears in behavior. And it is also difficult to assess them psychologically, because many of the feelings and judgments that we traditionally regard as positive may result from forms of censorship of other ideas, inhibition of competing activities, or suppression of more ambitious goals. It is possible that the importance of this subject itself tends rarely to be recognized, precisely because of the mass of inhibitory machinery that constitutes it. Could it be that our accumulations of counterexamples are larger and more powerful than our collections of instances and examples? Could it be that we learn more from negative rather than from positive reinforcement? Our hedonistic culture holds that learning works best when it seems pleasant and enjoyable -- but that discounts the value of experiencing frustrations, failures and disappointments, either in actuality or in the vicarious forms of forewarnings and admonishments.

References

Minsky, M. (1986) The Society of Mind, New York: Simon and Schuster

Minsky, M. (1991) "Society of Mind: A Response to Four Reviews." Artificial Intelligence, April 1991, Vol. 48, pages 371-396.

Agnew, N. Mck. & Brown, J.L. (1986). Bounded Rationality: Fallible decisions in unbounded decision space. Behavioral Science, 31, 148-161

Several paragraphs of this text are adapted from sections of "The Society of Mind."