Metacognitive Skill Learning in Humans and AI | Stanford
Smith Explains Dual System Metacognition
There are two basic forms of metacognitive information types: type one, which includes affective, non-propositional metadata and metacognitive feelings that are fast and automatic, such as feelings of knowing or rightness; and type two, which includes metacognitive strategies that are slower, declarative, and propositionally structured, such as concepts and strategies that allow us to direct our processes towards retrieving information. Type one metacognitive feelings can trigger type two metacognitive strategies, allowing us to retrieve knowledge that would otherwise remain inaccessible.
Smith on the Characteristics of Skill
Metacognitive skill embodies the same characteristics as motor and cognitive skill, including goal structure and knowledge types. Complex goals require sub-goals, entailing a hierarchical goal structure. Knowledge allows an agent to choose the right actions to achieve goals, directed by declarative knowledge (propositional facts, rules, explicit reasoning) and procedural knowledge (implicit representations that execute the control of the action). Procedural knowledge builds up over time through practice, becoming fast and automatic, eventually replacing declarative knowledge.
Smith Details Proceduralization Theory
The theory of metacognitive skill learning relies on a theory of proceduralization, where slow declarative knowledge is converted into fast procedural knowledge that is increasingly refined. Declarative knowledge moves into working memory, allowing you to activate actions, while procedural knowledge operates outside working memory. Over practice, procedural knowledge builds up to become faster, automatic, and eventually replace declarative knowledge, allowing for cognitive reinvestment where working memory is freed up for higher-level control. This allows you to apply your working memory to monitor the situation for changing circumstances, which allows you to apply your skills more flexibly.
Smith on Proceduralization Hallmarks
The hallmark signs of proceduralization seen in motor and cognitive skill should also be seen in metacognitive skill, specifically the power law function, which quantifies the speeding up of reaction times. Attentional skill and metamemory skill also follow a power law of learning, reflecting the characteristics of motor and cognitive skill. This theory also helps to understand confounding data in empirical research, such as why athletes who are regularly not very self-conscious are more likely to have their performance disrupted under pressure.
Smith on Metacognitive Proceduralization
Metacognitive proceduralization has been unidentified because it is invisible to observers and less perceivable to the performers themselves. Procedural knowledge is not conscious and operates outside of working memory, so when skill develops, it becomes an unconscious habit. As meta-knowledge is converted into imperceptible procedural knowledge, it disappears below the horizon, and people lose conscious access to the strategies they use to monitor and control their attention and learning strategies. This dual imperceptibility explains why metacognitive proceduralization has been unknown to date.
Smith on AI's Metacognitive Struggles
AI has been making incredible advancements, but it struggles the most in the metacognitive domain. There has been a trajectory towards increasing self-reference in AI design, with processes becoming increasingly self-monitoring, self-controlling, and self-learning. AI developers have noticed that using metacognition can significantly improve performance, such as in Metarag, which improves its accuracy by detecting and fixing reasoning errors, and Quietstar, which uses an inner monologue to consider the best option before outputting the answer.
Smith on Metalearning Reward Function
At triple AI, Smith proposed a metalearning reward function where metacognitive strategies would become automated as they are used and rewarded over time. The artificial system would identify a learning task, select the appropriate learning strategy, apply the strategy, and evaluate its effectiveness. If useful, it would result in a positive reward; if not, a negative reward. These rewards and punishments would feed back into its memory of strategies, where useful strategies would become more automated with time, allowing for more adaptive improvements under uncertainty.
