این مقاله روند تاریخی هوش مصنوعی را از بنیانهای نظری تا تحولات سال ۲۰۲۶ بررسی میکند.
تاریخ هوش مصنوعی؛ از رؤیای ماشین متفکر تا ظهور هوش مولد، مدلهای استدلالی و عاملهای هوشمند
هوش مصنوعی امروز چنان سریع وارد زندگی روزمره، اقتصاد، آموزش، پژوهش علمی، صنعت، رسانه و سیاست شده است که ممکن است این تصور ایجاد شود که با پدیدهای متعلق به همین چند سال اخیر روبهرو هستیم. ظهور ChatGPT در پایان سال ۲۰۲۲، گسترش مدلهای مولد تصویر و ویدئو، توسعه مدلهای چندوجهی و سپس ظهور سیستمهایی که میتوانند استدلال کنند، برنامه بنویسند، از ابزارهای مختلف استفاده کنند و بخشی از یک فرایند کاری را مستقلاً پیش ببرند، بدون تردید شتابی بیسابقه به این تحول داده است. با این همه، آنچه امروز «هوش مصنوعی» مینامیم حاصل بیش از هشت دهه تحول علمی است و ریشههای نظری آن حتی به سالهای پیش از پیدایش رایانههای الکترونیکی مدرن بازمیگردد. تاریخ هوش مصنوعی در واقع تاریخ تلاقی چند رشته است: منطق ریاضی، علوم اعصاب، نظریه محاسبه، آمار، مهندسی کامپیوتر، زبانشناسی، روانشناسی شناختی و، در دهههای اخیر، علوم داده و اقتصاد دیجیتال.
در مرکز این تاریخ یک پرسش بسیار قدیمی قرار دارد: آیا آنچه انسان «اندیشیدن» مینامد، دارای ساختاری است که بتوان آن را به مجموعهای از عملیات قابل محاسبه تبدیل کرد؟ اگر پاسخ مثبت باشد، آیا یک ماشین میتواند نه فقط اعداد را محاسبه کند، بلکه الگوها را تشخیص دهد، زبان را بفهمد، از تجربه بیاموزد، تصمیم بگیرد و حتی رفتارهایی از خود نشان دهد که ما آنها را هوشمندانه مینامیم؟
پاسخهایی که طی هشتاد سال گذشته به این پرسش داده شدهاند یکسان نبودهاند. در نخستین دوره تصور میشد هوش را میتوان به قواعد منطقی تبدیل کرد و آن قواعد را به ماشین داد. سپس معلوم شد که بسیاری از جنبههای هوش انسانی را نمیتوان بهآسانی در قالب قواعد صریح نوشت. در نتیجه، مرکز توجه از «برنامهریزی قواعد هوش» به «یادگیری از داده» منتقل شد. بعد شبکههای عصبی عمیق امکان استخراج الگوهای بسیار پیچیده را فراهم کردند. ظهور Transformer در سال ۲۰۱۷ راه را برای ساخت مدلهای بنیادین و مدلهای زبانی بسیار بزرگ گشود. سرانجام، از اوایل دهه ۲۰۲۰، هوش مصنوعی از مرحلهای که عمدتاً یک ابزار تخصصی بود به فناوریای عمومی تبدیل شد که قادر است با انسان از طریق زبان طبیعی ارتباط برقرار کند. تحول کنونی نیز آن را از «ماشینی که پاسخ میدهد» به سوی «سیستمی که میتواند برای رسیدن به یک هدف اقدام کند» سوق میدهد.
برای فهم این مسیر باید از زمانی آغاز کرد که هنوز اصطلاح «هوش مصنوعی» وجود نداشت. در دهه ۱۹۳۰، آلن تورینگ، ریاضیدان بریتانیایی، یکی از بنیادیترین پرسشهای علوم کامپیوتر را مطرح کرد: اصولاً چه چیزی قابل محاسبه است؟ مدل نظری او که بعدها «ماشین تورینگ» نام گرفت، نشان داد که میتوان مفهوم محاسبه را به شکلی دقیق و ریاضی تعریف کرد. اهمیت این دستاورد برای تاریخ هوش مصنوعی در این بود که مرزی نظری برای انجام عملیات توسط ماشین فراهم کرد. اگر فرایندی را بتوان به توالی مشخصی از عملیات تبدیل کرد، در اصل یک ماشین عمومی محاسبه نیز میتواند آن را اجرا کند.
گام مهم بعدی در سال ۱۹۴۳ برداشته شد؛ زمانی که Warren McCulloch، عصبشناس و روانپزشک، و Walter Pitts، منطقدان جوان، مقاله مشهور خود با عنوان «A Logical Calculus of the Ideas Immanent in Nervous Activity» را منتشر کردند. آنان مدلی انتزاعی از نورون ارائه کردند که در آن واحدهای ساده محاسباتی میتوانستند بسته به ورودیهایشان فعال یا غیرفعال شوند و در قالب شبکههایی به یکدیگر متصل گردند. این مدل با شبکههای عصبی امروزی تفاوتهای زیادی داشت، اما از نظر تاریخی بنیادی بود، زیرا برای نخستین بار ارتباطی رسمی میان عملکرد شبکه عصبی و منطق محاسباتی برقرار میکرد. به بیان دیگر، این تصور شکل گرفت که شاید بتوان برخی عملکردهای سیستم عصبی را با شبکهای از عناصر محاسباتی مدلسازی کرد. مقاله McCulloch و Pitts که در Bulletin of Mathematical Biophysics منتشر شد، بعدها یکی از پایههای نظری شبکههای عصبی مصنوعی به شمار آمد.
در سال ۱۹۵۰ تورینگ پرسش را یک مرحله جلوتر برد. مقاله او با عنوان «Computing Machinery and Intelligence» در مجله Mind با این پرسش آغاز شد که آیا ماشینها میتوانند فکر کنند. تورینگ به جای گرفتار شدن در تعریف فلسفی واژههای «ماشین» و «تفکر»، پیشنهاد کرد مسئله را از طریق رفتار بررسی کنیم. او «بازی تقلید» را مطرح کرد؛ آزمایشی که بعدها به آزمون تورینگ شهرت یافت. در سادهترین تعبیر آن، اگر یک انسان از طریق گفتوگوی متنی نتواند با اطمینان تشخیص دهد که طرف مقابل انسان است یا ماشین، ماشین رفتاری از خود نشان داده که از نظر عملی میتوان آن را هوشمندانه تلقی کرد. اهمیت مقاله تورینگ فقط در ارائه یک آزمون نبود؛ او بسیاری از استدلالهایی را که هنوز هم درباره امکان هوش ماشینی مطرح میشوند بررسی کرد و نشان داد که مسئله رابطه میان محاسبه و هوش باید به یک مسئله پژوهشی جدی تبدیل شود.
با این همه، اصطلاح «Artificial Intelligence» هنوز وجود نداشت. نقطهای که معمولاً تولد رسمی هوش مصنوعی بهعنوان یک رشته علمی تلقی میشود، پروژه تابستانی Dartmouth در سال ۱۹۵۶ است. پیشنهاد اولیه این پروژه در سال ۱۹۵۵ توسط John McCarthy، Marvin Minsky، Nathaniel Rochester و Claude Shannon تهیه شده بود. McCarthy اصطلاح «Artificial Intelligence» را برای نامیدن حوزه جدید به کار برد. فرض اساسی پیشنهاد Dartmouth بسیار بلندپروازانه بود: این پژوهشگران تصور میکردند هر جنبهای از یادگیری یا دیگر ویژگیهای هوش را میتوان در اصل آنقدر دقیق توصیف کرد که ماشین قادر به شبیهسازی آن باشد. گردهمایی تابستان ۱۹۵۶ در Dartmouth گروهی از افرادی را کنار هم آورد که بعدها به بنیانگذاران اصلی رشته AI تبدیل شدند و از همین رو سال ۱۹۵۶ معمولاً تاریخ تولد رسمی این رشته محسوب میشود.
دهههای ۱۹۵۰ و ۱۹۶۰ دوران خوشبینی فوقالعاده بود. رایانهها تازه در حال شکلگیری بودند و نخستین موفقیتها این تصور را تقویت میکردند که شاید ایجاد ماشینهای واقعاً هوشمند چندان دور نباشد. یکی از نخستین برنامههای مشهور، Logic Theorist بود که Allen Newell، Herbert Simon و Cliff Shaw آن را توسعه دادند. این برنامه میتوانست برخی قضایای منطق ریاضی را اثبات کند. کمی بعد General Problem Solver ساخته شد و تلاش کرد روشهایی عمومی برای حل مسائل ارائه دهد. در این دوران رویکردی شکل گرفت که بعدها «هوش مصنوعی نمادین» یا Symbolic AI نامیده شد. فرض آن این بود که جهان را میتوان در قالب اشیا، مفاهیم و روابط نمادین توصیف کرد و استدلال را نیز به قواعد منطقی تبدیل نمود. اگر دانش کافی درباره یک حوزه به ماشین داده شود و قواعد استنتاج نیز مشخص باشند، ماشین قادر خواهد بود نتیجهگیری کند.
این رویکرد برای مسائل محدود بسیار موفق بود، اما محدودیتی اساسی داشت: انسانها بسیاری از کارهای روزمره خود را بدون داشتن مجموعهای صریح از قواعد انجام میدهند. برای نمونه، تشخیص چهره یک دوست در خیابان، فهم یک شوخی، تشخیص لحن گوینده یا تمایز دادن میان هزاران شیء متفاوت بهسادگی قابل تبدیل به هزاران جمله «اگر… آنگاه…» نیست. هوش انسانی بخش عظیمی از توان خود را از تجربه و یادگیری الگوهایی به دست میآورد که انسان حتی نمیتواند همیشه آنها را بهصورت قواعد صریح بیان کند.
همزمان با رویکرد نمادین، مسیر دیگری نیز در حال شکلگیری بود که بعدها نقشی تعیینکننده یافت. Frank Rosenblatt در اواخر دهه ۱۹۵۰ مدل Perceptron را توسعه داد؛ سامانهای الهامگرفته از نورون که میتوانست بر اساس نمونههای آموزشی وزنهای ارتباطی خود را تغییر دهد. اهمیت پرسپترون از این جهت بود که به جای آنکه برنامهنویس تمام قواعد تصمیمگیری را از پیش تعیین کند، سیستم میتوانست برخی روابط را از دادههای نمونه یاد بگیرد. این همان تفاوت بنیادی است که بعدها میان برنامهنویسی سنتی و یادگیری ماشین شکل گرفت: به جای اینکه انسان دقیقاً به رایانه بگوید برای هر ورودی چه کند، نمونههایی در اختیار آن گذاشته میشود تا خود الگویی برای تصمیمگیری پیدا کند. پژوهش Rosenblatt در سال ۱۹۵۸ در Psychological Review منتشر شد و یکی از اجداد مستقیم شبکههای عصبی امروزی به شمار میرود.
اما محدودیتهای سختافزاری و نظری خیلی زود خود را نشان دادند. رایانهها قدرت پردازش اندکی داشتند، حافظه بسیار محدود بود و حجم داده دیجیتال با امروز قابل مقایسه نبود. مهمتر از همه، پژوهشگران دریافتند مسائلی که برای انسان بسیار ساده به نظر میرسند میتوانند برای ماشین فوقالعاده دشوار باشند. انسان کودک خردسالی را با چند نمونه قادر میکند گربه را از سگ تشخیص دهد، در حالی که آموزش ماشین برای همان کار به حجم عظیمی از داده نیاز داشت. ترجمه زبان، درک گفتار، تشخیص اشیا و استدلال درباره موقعیتهای جهان واقعی بسیار دشوارتر از اثبات قضایای محدود منطقی بودند.
در دهه ۱۹۷۰ فاصله میان وعدههای اولیه و دستاوردهای عملی باعث کاهش حمایت مالی در برخی مراکز شد. این دوره بعدها «زمستان هوش مصنوعی» نام گرفت. اصطلاح AI Winter به دورههایی اشاره دارد که پس از موجهای خوشبینی و سرمایهگذاری، عدم تحقق انتظارات موجب کاهش بودجه و توجه عمومی شد. هوش مصنوعی در تاریخ خود بیش از یک چنین دورهای را تجربه کرد و همین فراز و فرودها نشان میدهد که پیشرفت AI نه خطی بوده و نه اجتنابناپذیر؛ هر جهش آن به مجموعهای از پیشرفتهای علمی، سختافزاری و اقتصادی نیاز داشته است.
در دهه ۱۹۸۰ هوش مصنوعی بار دیگر، این بار با «سیستمهای خبره» یا Expert Systems، مورد توجه قرار گرفت. در این سیستمها دانش متخصصان یک حوزه در قالب مجموعه بزرگی از قواعد ذخیره میشد. برای مثال یک سیستم پزشکی میتوانست بر اساس ترکیبی از علائم و نتایج آزمایش احتمالاتی درباره بیماری ارائه دهد یا یک سیستم صنعتی برای پیکربندی تجهیزات پیشنهادهایی مطرح کند. MYCIN در پزشکی و XCON در صنعت کامپیوتر نمونههای شناختهشده این رویکرد بودند. سیستمهای خبره نشان دادند AI میتواند در حوزههای محدود اقتصادی ارزش عملی داشته باشد، اما مشکل نگهداری هزاران قاعده به تدریج آشکار شد. «مهندسی دانش» به فرایندی پرهزینه تبدیل میشد: دانش متخصص باید استخراج، صریح، ثبت و دائماً بهروز میشد. جهان واقعی بیش از اندازه پیچیده و پویا بود که همیشه بتوان آن را در مجموعهای از قواعد ثابت محصور کرد.
یکی از تحولات علمی مهم همین دهه، احیای شبکههای عصبی بود. در سال ۱۹۸۶ David Rumelhart، Geoffrey Hinton و Ronald Williams مقاله مشهور «Learning representations by back-propagating errors» را در Nature منتشر کردند. این مقاله استفاده مؤثر از الگوریتم backpropagation را برای تنظیم وزنهای شبکههای چندلایه توضیح میداد. اصل کار آن است که شبکه ابتدا خروجی تولید میکند، اختلاف میان خروجی واقعی و خروجی مورد انتظار اندازهگیری میشود و سپس خطا از لایههای پایانی به عقب منتقل میگردد تا وزن ارتباطات به نحوی تغییر کنند که خطا کاهش یابد. این روش بعدها به یکی از ستونهای اصلی آموزش شبکههای عصبی عمیق تبدیل شد.
با این همه، تا مدتها شبکههای عصبی محدود باقی ماندند. برای آنکه توان واقعی آنها آشکار شود سه عامل باید همزمان رشد میکردند: داده، قدرت محاسباتی و الگوریتم. این سه عامل از دهه ۱۹۹۰ و بهویژه در دهه ۲۰۰۰ به تدریج به یکدیگر رسیدند. اینترنت، تلفنهای هوشمند، تجارت الکترونیک، شبکههای اجتماعی، سیستمهای دیجیتال و حسگرها حجم عظیمی از داده ایجاد کردند. توان پردازندهها چندین مرتبه افزایش یافت و GPUها، که در اصل برای پردازش گرافیک توسعه یافته بودند، مشخص شد برای عملیات موازی مورد نیاز شبکههای عصبی بسیار مناسباند. همزمان الگوریتمها، روشهای تنظیم شبکه، مجموعهدادههای بزرگ و ابزارهای نرمافزاری بهبود یافتند.
در این دوره معنای هوش مصنوعی نیز آرامآرام تغییر کرد. به جای آنکه پژوهشگر همه دانش را به ماشین بدهد، ماشین با استفاده از دادههای بزرگ به استخراج روابط آماری میپرداخت. این همان حوزهای است که Machine Learning یا یادگیری ماشین نام گرفت. برای مثال، در روش سنتی تشخیص ایمیل هرزنامه، برنامهنویس ممکن بود دهها قاعده تعریف کند: اگر این کلمات وجود داشتند یا فرستنده چنین ویژگیهایی داشت احتمال Spam بیشتر است. در یادگیری ماشین، هزاران یا میلیونها نمونه ایمیل با برچسب «هرزنامه» یا «عادی» در اختیار الگوریتم قرار میگیرد و مدل خودش ترکیبی از ویژگیهای مؤثر را یاد میگیرد. بنابراین تحول مهمی رخ داد: دانش ماشین بهتدریج از چیزی که مستقیماً توسط برنامهنویس نوشته میشد، به چیزی تبدیل شد که از داده استخراج میشد.
یکی از نمادهای دوره میانی AI در سال ۱۹۹۷ رقم خورد، زمانی که سیستم Deep Blue شرکت IBM توانست Garry Kasparov، قهرمان جهان شطرنج، را در یک مسابقه رسمی شکست دهد. در افکار عمومی این رویداد بسیار مهم بود، زیرا شطرنج قرنها نمادی از تفکر استراتژیک و هوش انسانی محسوب میشد. با این حال Deep Blue با مدلهای AI امروز تفاوت بنیادی داشت. این سیستم برای حوزه خاص شطرنج ساخته شده بود و از جستوجوی گسترده، ارزیابی موقعیتها و دانش تخصصی استفاده میکرد. شکست Kasparov نشان داد ماشین میتواند در یک حوزه محدود از انسان پیشی بگیرد، اما به معنای وجود هوشی عمومی و انعطافپذیر نبود.
نقطه عطف بزرگ بعدی در سال ۲۰۱۲ پدید آمد. Alex Krizhevsky، Ilya Sutskever و Geoffrey Hinton شبکهای عمیق را برای مسابقه تشخیص تصویر ImageNet آموزش دادند. شبکهای که بعدها AlexNet نامیده شد، روی بیش از یک میلیون تصویر آموزش دید و توانست میزان خطا را به شکل چشمگیری نسبت به روشهای پیشین کاهش دهد. مقاله آنان در کنفرانس NeurIPS نشان داد شبکههای عصبی عمیق، در صورتی که داده و قدرت محاسباتی کافی در اختیار داشته باشند، قادرند ویژگیهای پیچیده تصویر را خودشان یاد بگیرند. این رویداد در عمل آغاز انفجار Deep Learning در دهه ۲۰۱۰ بود. پس از آن شبکههای عمیق به سرعت در تشخیص تصویر، گفتار، ترجمه، پردازش زبان، پزشکی و بسیاری حوزههای دیگر به روش غالب تبدیل شدند.
Deep Learning یا یادگیری عمیق اساساً شبکههای عصبی چندلایهای است که هر لایه قادر است نمایش پیچیدهتری از داده ایجاد کند. برای نمونه، در یک سیستم تشخیص تصویر ممکن است لایههای اولیه الگوهای سادهای مانند لبهها را تشخیص دهند، لایههای میانی ترکیبهایی مانند اشکال یا بافتها را بیاموزند و لایههای عمیقتر ساختارهایی مانند چشم، صورت، حیوان یا خودرو را تشخیص دهند. ویژگی مهم این تحول آن بود که مهندس دیگر لازم نبود تمام ویژگیهای لازم برای تشخیص تصویر را دستی تعریف کند؛ شبکه آنها را از داده استخراج میکرد.
چهار سال پس از AlexNet، رویداد نمادین دیگری تصور عمومی از AI را تغییر داد. در مارس ۲۰۱۶ AlphaGo، سیستم شرکت DeepMind، در یک مسابقه پنجبازی Lee Sedol، یکی از برجستهترین بازیکنان جهان در بازی Go را با نتیجه چهار بر یک شکست داد. Go به دلیل تعداد نجومی حالتهای ممکن بسیار پیچیدهتر از شطرنج است و مدتها تصور میشد ماشین برای رسیدن به سطح قهرمانان انسانی در آن به زمان بسیار بیشتری نیاز دارد. AlphaGo از شبکههای عصبی، یادگیری تقویتی و جستوجوی درختی استفاده میکرد. پیروزی آن نشان داد ترکیب یادگیری و محاسبات میتواند به راهحلهایی برسد که به مراتب فراتر از جستوجوی خام تمام حالتهای ممکن است.
نسخههای بعدی این ایده حتی فراتر رفتند. AlphaGo Zero و AlphaZero نشان دادند که سیستم میتواند با بازی کردن علیه خودش و از طریق Reinforcement Learning، بدون نیاز به انبوه بازیهای انسانی، استراتژیهای بسیار قدرتمندی بیاموزد. مقاله AlphaGo Zero در Nature توضیح داد که سیستم میتوانست تنها بر اساس قواعد بازی و خودبازی، شبکه خود را بهبود دهد و عملاً «معلم خودش» شود. اینجا مفهوم یادگیری ماشین به مرحله تازهای رسید: ماشین نه فقط از دادههای تولیدشده توسط انسان، بلکه از تجربهای که خودش ایجاد میکرد نیز میآموخت.
با این همه، انقلابی که مستقیماً به مدلهای زبانی بزرگ امروز انجامید، در سال ۲۰۱۷ رخ داد. گروهی از پژوهشگران Google مقالهای با عنوان «Attention Is All You Need» منتشر کردند و معماری Transformer را معرفی نمودند. تا آن زمان بسیاری از سیستمهای پردازش زبان از شبکههای بازگشتی یا روشهایی استفاده میکردند که پردازش توالیهای طولانی را دشوار میساخت. Transformer بر سازوکاری به نام Attention تکیه کرد که به مدل اجازه میدهد هنگام پردازش هر بخش از متن، میزان ارتباط آن را با بخشهای دیگر ارزیابی کند. مزیت مهم دیگر این بود که Transformer امکان موازیسازی بسیار بیشتری در جریان آموزش فراهم میکرد. نویسندگان مقاله نشان دادند این معماری در ترجمه ماشینی عملکردی بهتر و آموزش سریعتری نسبت به بسیاری از روشهای قبلی دارد. اهمیت تاریخی این مقاله بعدها آشکار شد: معماری Transformer بنیان اصلی بسیاری از مدلهای زبان بزرگ، مدلهای چندوجهی و سیستمهای مولد امروزی شد.
پس از Transformer مفهوم «مدل بنیادین» یا Foundation Model اهمیت فزایندهای یافت. به جای ساخت یک مدل کوچک جداگانه برای ترجمه، یکی دیگر برای خلاصهسازی و مدل دیگری برای پاسخ به سؤال، میتوان یک مدل بسیار بزرگ را روی حجم عظیمی از اطلاعات آموزش داد و سپس از همان مدل برای وظایف متعدد استفاده کرد. مدلهای GPT یکی از مهمترین مسیرهای این تحول بودند. GPT-2 در سال ۲۰۱۹ نشان داد که مدل زبانی بزرگ آموزشدیده روی متون وسیع میتواند متنهای نسبتاً منسجم تولید کند و بدون آموزش جداگانه در وظایفی مانند پاسخ به سؤال، خلاصهسازی و ترجمه قابلیتهایی از خود نشان دهد.
در سال ۲۰۲۰ GPT-3 با ۱۷۵ میلیارد پارامتر جهش دیگری ایجاد کرد. مقاله «Language Models are Few-Shot Learners» نشان داد که صرف افزایش مقیاس مدل میتواند توانایی آن را برای انجام وظایف متعدد از طریق نمونههای اندک یا حتی دستورهای متنی افزایش دهد. این تحول از لحاظ رابط انسان و ماشین بسیار مهم بود. در برنامهنویسی سنتی، برای تغییر رفتار نرمافزار باید کد تغییر کند؛ اما در مدلهای زبانی بزرگ، کاربر به تدریج میتوانست تنها از طریق زبان طبیعی رفتار سیستم را هدایت کند. Prompt به نوعی رابط جدید میان خواست انسان و فرایند محاسباتی تبدیل شد.
در اینجا باید به یک تحول فنی و اجتماعی دیگر نیز توجه کرد: آموزش مدل برای «پاسخ مطلوب به انسان». پیشبینی توکن بعدی به خودی خود تضمین نمیکند که مدل سؤال کاربر را درست بفهمد یا پاسخ مفیدی بدهد. توسعه روشهایی مانند Reinforcement Learning from Human Feedback یا RLHF به مدل اجازه داد از ترجیحات انسانی درباره کیفیت پاسخها نیز یاد بگیرد. InstructGPT نمونه مهم این مسیر بود و تحقیقات منتشرشده در سال ۲۰۲۲ نشان دادند که تنظیم مدل با بازخورد انسانی میتواند میزان پیروی از دستورها و کیفیت تعامل را افزایش دهد.
سپس در ۳۰ نوامبر ۲۰۲۲ ChatGPT عرضه شد. از نظر علمی، بسیاری از اجزای اصلی آن از قبل وجود داشتند، اما از نظر اجتماعی ChatGPT یک نقطه عطف بود. رابط گفتوگویی پیچیدگی تعامل با مدلهای بزرگ را از دید کاربر پنهان کرد. اکنون فردی که هیچ دانش برنامهنویسی نداشت میتوانست با زبان روزمره از ماشین بخواهد مقالهای را خلاصه کند، برنامهای بنویسد، یک مفهوم علمی را توضیح دهد، متنی را ویرایش کند، ایدهپردازی کند یا دادهای را تحلیل نماید. OpenAI در معرفی اولیه ChatGPT در نوامبر ۲۰۲۲ آن را مدلی گفتوگومحور مرتبط با InstructGPT معرفی کرد که برای پیروی از دستورهای کاربر طراحی شده بود.
از این نقطه، AI وارد مرحلهای شد که میتوان آن را «دموکراتیزه شدن رابط هوش مصنوعی» نامید. تا پیش از آن، استفاده مؤثر از بسیاری از سیستمهای AI نیازمند برنامهنویسی یا مهارت تخصصی بود. مدلهای مکالمهای زبان طبیعی را به رابطی عمومی تبدیل کردند. این امر اهمیت تاریخی دارد، زیرا زبان طبیعی ابزار اصلی ارتباط اجتماعی انسان است. اگر رایانههای شخصی رابط گرافیکی را به رابط عمومی انسان و کامپیوتر تبدیل کردند، مدلهای زبانی بزرگ امکان آن را ایجاد کردهاند که زبان روزمره خود انسان به یکی از رابطهای اصلی کنترل سیستمهای دیجیتال تبدیل شود.
اما تحول در مدلهای زبان متوقف نماند. مدلها به تدریج «چندوجهی» یا Multimodal شدند. انسان جهان را فقط از طریق متن تجربه نمیکند؛ تصویر، صدا، زبان، حرکت و محیط فیزیکی همگی بخشی از شناخت هستند. مدلهای چندوجهی نیز میتوانند اطلاعات را از چند قالب دریافت یا تولید کنند. یک سیستم میتواند تصویر را تحلیل کند و درباره آن صحبت کند، یک نمودار را توضیح دهد، صدای انسان را دریافت کند، با او مکالمه کند، تصاویر تولید کند یا اطلاعات متنی و تصویری را با یکدیگر ترکیب نماید. به این ترتیب مرز سنتی میان «سیستم تشخیص تصویر»، «مدل زبان»، «سیستم صوتی» و «نرمافزار» در حال کمرنگ شدن است.
تحول مهم بعدی، برجسته شدن مدلهای استدلالی بود. مدلهای زبانی اولیه در تولید متن روان بسیار قوی بودند، اما در مسائل پیچیده چندمرحلهای، ریاضیات یا برنامهنویسی اغلب دچار خطا میشدند. نسلهای جدیدتر AI بیشتر بر فرایند حل مسئله متمرکز شدهاند: تجزیه یک هدف پیچیده به مراحل مختلف، بررسی راهحلهای متفاوت، استفاده از ابزارهای بیرونی، اجرای محاسبات و ارزیابی نتیجه. گزارش AI Index 2026 دانشگاه Stanford نشان میدهد که پیشرفت در برخی آزمونهای دشوار علمی، ریاضی، استدلال چندوجهی و برنامهنویسی در فاصله کوتاهی بسیار سریع بوده است. طبق این گزارش، برخی مدلهای مرزی در شماری از آزمونها به معیارهای انسانی رسیده یا از آنها عبور کردهاند و عملکرد روی بعضی بنچمارکهای برنامهنویسی نیز در یک سال افزایش بسیار بزرگی داشته است.
در کنار مدلهای استدلالی، مفهوم Agentic AI یا «هوش مصنوعی عاملمحور» به یکی از مهمترین جهتهای توسعه تبدیل شده است. تفاوت میان chatbot و agent اساسی است. یک chatbot معمولاً سؤال دریافت میکند و پاسخ میدهد. اما یک عامل هوشمند میتواند هدفی دریافت کند، آن را به مجموعهای از مراحل تبدیل کند، درباره ترتیب کار تصمیم بگیرد، از موتور جستوجو، پایگاه داده، نرمافزار، فایل یا کد استفاده کند، نتیجه یک مرحله را بررسی کند و بر اساس آن مرحله بعدی را انتخاب نماید. در چنین مدلی AI از یک «ماشین پاسخدهنده» به نوعی «سامانه انجام وظیفه» تبدیل میشود. پژوهشگران Dartmouth نیز در سال ۲۰۲۶، ضمن بررسی Agentic AI، بر این نکته تأکید کردهاند که سیستمهای عاملمحور مرحله مهمی در توسعه AI هستند، هرچند همچنان با محدودیتهایی در قابلیت اعتماد و خودمختاری روبهرو هستند.
این تحول میتواند از نظر اقتصادی حتی عمیقتر از ظهور chatbotها باشد. هنگامی که AI فقط یک متن تولید میکند، ابزار افزایش بهرهوری است؛ اما هنگامی که میتواند بخشی از یک فرایند تولید، تحقیق، برنامهنویسی، مدیریت، پشتیبانی یا تحلیل را از ابتدا تا انتها پیش ببرد، به یک جزء فعال در سازماندهی کار تبدیل میشود. بنابراین مسئله آینده AI تنها این نیست که ماشین «چه میداند»، بلکه این است که چه میزان از چرخه تبدیل هدف به تصمیم و تصمیم به عمل میتواند به سیستمهای هوشمند واگذار شود.
تا سال ۲۰۲۶، هوش مصنوعی به مرحلهای رسیده است که توصیف آن صرفاً به عنوان یک رشته علوم کامپیوتر دیگر کافی نیست. این فناوری در حال تبدیل شدن به زیرساخت عمومی فعالیتهای شناختی است. مدلهای پیشرفته میتوانند متن، تصویر، صدا و کد را پردازش کنند، در تحقیقات علمی کمک کنند، دادههای بزرگ را تحلیل کنند، در تولید نرمافزار مشارکت کنند، مسائل ریاضی حل کنند، با ابزارهای خارجی ارتباط برقرار نمایند و در قالب agent وظایف چندمرحلهای را دنبال کنند. Stanford AI Index 2026 همچنین گزارش میکند که بیش از ۹۰ درصد مدلهای برجسته مرزی سال ۲۰۲۵ توسط صنعت تولید شدهاند. این واقعیت نشان میدهد مرکز ثقل پژوهش AI که زمانی عمدتاً در دانشگاهها و آزمایشگاههای عمومی قرار داشت، در حوزه مدلهای مرزی به شرکتهایی منتقل شده که توانایی تأمین هزینه عظیم تراشه، مراکز داده، برق، شبکه و نیروی متخصص را دارند.
در اینجا تاریخ تکنیکی هوش مصنوعی به تاریخ اقتصادی آن پیوند میخورد. AI جدید صرفاً یک الگوریتم نیست؛ حاصل ترکیب الگوریتم، داده، تراشه، انرژی، مراکز داده، شبکه ارتباطی و سرمایه عظیم است. به همین دلیل بحث درباره هوش مصنوعی امروز ناگزیر بحث درباره مالکیت زیرساختهای آن نیز هست. یک مدل پیشرفته ممکن است بر حجم عظیمی از دانش انباشته انسانی ــ کتابها، مقالات علمی، نرمافزارها، تصاویر، زبان و دیگر آثار فرهنگی ــ آموزش دیده باشد، اما توانایی ساخت و کنترل مدلهای مرزی در اختیار شمار نسبتاً محدودی از نهادهای بزرگ اقتصادی قرار دارد. در این نقطه یک تضاد قابل توجه تاریخی شکل میگیرد: مواد اولیه معرفتی AI عمیقاً اجتماعی هستند، در حالی که ابزار پردازش و تبدیل آنها به قدرت اقتصادی میتواند بسیار متمرکز باشد.
همین مسئله به یکی از بزرگترین پرسشهای اجتماعی عصر هوش مصنوعی منتهی میشود. پرسش آینده فقط این نیست که «ماشین چه کارهایی میتواند انجام دهد؟» بلکه باید پرسید چه کسی مالک ماشین، داده، مدل، تراشه و زیرساخت است؛ چه کسی اهداف سیستم را تعیین میکند؛ چه کسی حق دسترسی دارد؛ چه کسانی از افزایش بهرهوری آن بهرهمند میشوند؛ و هزینههای اجتماعی آن بر دوش چه کسانی قرار میگیرد.
تاریخ هوش مصنوعی از این زاویه بخشی از تاریخ نیروهای مولد نیز هست. ماشین بخار قدرت عضلانی انسان را چند برابر کرد. موتور الکتریکی و خط تولید امکان سازماندهی تولید انبوه را فراهم کردند. رایانه محاسبه و پردازش اطلاعات را خودکار کرد. اینترنت امکان انتقال و اشتراک تقریباً آنی اطلاعات را در مقیاس جهانی ایجاد نمود. هوش مصنوعی اکنون وارد قلمروی دیگری شده است: بخشی از کار شناختی.
کار شناختی شامل نوشتن، ترجمه، تحلیل، طراحی، برنامهنویسی، تشخیص الگو، برنامهریزی، تحقیق و بخشی از تصمیمسازی است؛ فعالیتهایی که تا همین اواخر تصور میشدند بهطور انحصاری یا تقریباً انحصاری به انسان تعلق دارند. هوش مصنوعی الزاماً تمام این فعالیتها را جایگزین نمیکند، اما رابطه میان انسان و ابزار تولید را در آنها تغییر میدهد. همانگونه که ورود ماشین صنعتی الزاماً کار فیزیکی انسان را فوراً حذف نکرد اما ماهیت و سازمان آن را دگرگون کرد، AI نیز در حال تغییر ساختار کار فکری است.
با این حال باید از یک اشتباه مهم پرهیز کرد: هوش مصنوعی و آگاهی یک مفهوم نیستند. پیشرفت حیرتآور مدلهای امروز نشان میدهد ماشین میتواند در برخی وظایف عملکردی را تولید کند که پیشتر نشانه هوش انسانی دانسته میشد، اما این امر به خودی خود نشان نمیدهد که ماشین دارای consciousness یا تجربه ذهنی است. یک مدل میتواند هزاران صفحه درباره درد، عشق، ترس یا مرگ تولید کند بدون آنکه شواهدی داشته باشیم که خود آن تجربیات را احساس میکند. میتواند درباره یک مفهوم استدلال کند بدون آنکه بتوانیم از این عملکرد نتیجه بگیریم دارای تجربه خودآگاهانهای مشابه انسان است. بنابراین باید میان پردازش اطلاعات، یادگیری، هوش، شناخت و آگاهی تمایز قائل شد.
این تمایز از نظر آینده اجتماعی AI اهمیت زیادی دارد. ممکن است جامعهای دارای بالاترین سطح هوش محاسباتی باشد، اما الزاماً جامعهای آگاهتر نباشد. یک سیستم فوقالعاده قدرتمند میتواند برای کشف دارو، آموزش، برنامهریزی اقتصادی، کاهش اتلاف منابع یا گسترش دسترسی به دانش مورد استفاده قرار گیرد؛ همان سیستم یا فناوری مشابه میتواند در نظارت گسترده، دستکاری اطلاعات، جنگ، انحصار اقتصادی یا کنترل اجتماعی نیز به کار رود. فناوری به خودی خود جهت اخلاقی یا اجتماعی نهایی خود را تعیین نمیکند. این انسانها، نهادها و ساختارهای اجتماعیاند که تعیین میکنند ظرفیت فنی چگونه مورد استفاده قرار گیرد.
از همین رو، تاریخ هوش مصنوعی را میتوان در دو سطح خواند. در سطح نخست، یک تاریخ تکنیکی مشاهده میکنیم: منطق ریاضی به محاسبات ماشینی، محاسبات به هوش نمادین، هوش نمادین به یادگیری ماشین، یادگیری ماشین به شبکههای عمیق، شبکههای عمیق به Transformer، Transformer به مدلهای بنیادین، مدلهای بنیادین به هوش مولد و چندوجهی، و آنها به مدلهای استدلالی و عاملهای هوشمند منتهی شدهاند.
اما در سطح عمیقتر، تحول دیگری رخ داده است: ابتدا انسان قواعد را به ماشین میداد؛ سپس داده در اختیار آن گذاشت تا الگو را یاد بگیرد؛ سپس ماشین توانست بازنماییهای پیچیدهای از اطلاعات ایجاد کند؛ مدلهای بزرگ بخشی از دانش اجتماعی انباشتهشده را در ساختارهای آماری خود جذب کردند؛ و اکنون سیستمهای عاملمحور میکوشند این دانش را به اقدام تبدیل کنند.
به بیان فشرده، مسیر تحول را میتوان چنین دید:
قاعده → داده → اطلاعات → الگو → مدل → دانش عملیاتی → استدلال → اقدام
این زنجیره شاید از هر جدول زمانی دیگری بهتر توضیح دهد که چرا هوش مصنوعی امروز با فناوریهای پیشین تفاوت دارد. ابزار دیگر صرفاً چیزی نیست که انسان از آن برای اجرای یک دستور ثابت استفاده کند؛ ابزار به تدریج در تفسیر هدف، یافتن مسیر و انتخاب بعضی اقدامات نیز مشارکت میکند.
در نتیجه، پرسشی که آلن تورینگ در سال ۱۹۵۰ مطرح کرد ــ آیا ماشینها میتوانند فکر کنند؟ ــ هنوز از بین نرفته است، اما در سال ۲۰۲۶ دیگر تنها پرسش مهم نیست. ماشینها اکنون قادرند رفتارهایی تولید کنند که در بسیاری حوزهها عملاً هوشمندانه تلقی میشوند. مسئله اجتماعی بزرگتر این است که انسان با این ظرفیت جدید چه خواهد کرد؟
آیا هوش مصنوعی عمدتاً ابزاری برای تمرکز بیشتر ثروت و قدرت خواهد شد، یا میتواند به گسترش دانش و مشارکت اجتماعی کمک کند؟ آیا بهرهوری حاصل از آن به کاهش ساعات کار و ارتقای کیفیت زندگی منجر خواهد شد، یا به حذف فرصتهای شغلی بدون توزیع منافع؟ آیا داده و دانش اجتماعی به منبع عمومی توانمندسازی انسان تبدیل خواهند شد یا ماده اولیه انحصارات جدید؟ آیا AI ابزار شفافیت قدرت خواهد بود یا ابزار نظارت بر شهروند؟ و مهمتر از همه، آیا افزایش «هوش» ماشینی به افزایش «آگاهی» انسانی و اجتماعی خواهد انجامید؟
شاید همینجا باشد که تاریخ هوش مصنوعی با بحث بزرگتری درباره نظم اجتماعی نوین در عصر آگاهی پیوند میخورد. انقلاب صنعتی پرسش مالکیت کارخانه و ابزار تولید را به مرکز تحولات اجتماعی آورد. انقلاب دیجیتال مسئله مالکیت اطلاعات و شبکهها را برجسته کرد. انقلاب هوش مصنوعی اکنون مسئله مالکیت و کنترل زیرساختهایی را مطرح میکند که نه فقط کالاهای مادی، بلکه دانش، تصمیم و فعالیت شناختی را پردازش میکنند.
از این منظر، AI را نمیتوان فقط بهعنوان یک اختراع فنی بررسی کرد. این فناوری در حال تبدیل شدن به یکی از نیروهای مولد تعیینکننده قرن بیستویکم است. همانگونه که فناوریهای گذشته ساختار اقتصاد و جامعه را تغییر دادند، هوش مصنوعی نیز احتمالاً ساختار کار، آموزش، مالکیت، قدرت، سیاست و حتی تعریف ما از مهارت انسانی را دگرگون خواهد کرد.
بنابراین مهمترین فصل تاریخ هوش مصنوعی شاید هنوز نوشته نشده باشد. نخستین فصل آن درباره این بود که آیا میتوان ماشین را وادار به محاسبه کرد. فصل بعد پرسید آیا ماشین میتواند استدلال کند. سپس مسئله یادگیری، دیدن، شنیدن و فهم زبان مطرح شد. امروز سؤال به سمت توانایی ماشین برای حل مسئله و عمل نسبتاً مستقل حرکت کرده است. اما فصل آینده بیش از آنکه مسئلهای درباره ماشین باشد، مسئلهای درباره جامعه خواهد بود.
تا امروز پرسش اصلی دانشمندان این بوده است:
ماشین تا چه اندازه میتواند هوشمند شود؟
پرسش تعیینکننده عصر آینده احتمالاً این خواهد بود:
جامعه انسانی با این هوش چه خواهد کرد؟
پاسخ به پرسش نخست عمدتاً در آزمایشگاهها، دانشگاهها و مراکز داده نوشته شده است. پاسخ به پرسش دوم را دیگر مهندسان به تنهایی نمیتوانند تعیین کنند. این پاسخ به اقتصاد، سیاست، اخلاق، حقوق، فرهنگ، مالکیت، مشارکت اجتماعی و در نهایت به سطح آگاهی جامعه وابسته خواهد بود.
در همین معنا، تاریخ هوش مصنوعی دیگر فقط تاریخ ماشینهای هوشمند نیست؛ به تدریج به بخشی از تاریخ تحول خود جامعه انسانی تبدیل شده است.
منابع و مطالعات بیشتر
- Turing, A. M. (1950). Computing Machinery and Intelligence. Mind, 59(236), 433–460. مقاله بنیادی تورینگ که مسئله «آیا ماشینها میتوانند فکر کنند؟» و بازی تقلید را مطرح کرد.
- McCulloch, W. S., & Pitts, W. (1943). A Logical Calculus of the Ideas Immanent in Nervous Activity. Bulletin of Mathematical Biophysics, 5, 115–133. یکی از نخستین مدلهای رسمی نورون و شبکه عصبی مصنوعی.
- McCarthy, J., Minsky, M., Rochester, N., & Shannon, C. (1955). A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence. پیشنهاد تاریخی پروژهای که در تابستان ۱۹۵۶ به شکلگیری رسمی رشته هوش مصنوعی انجامید.
- Rosenblatt, F. (1958). The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain. Psychological Review, 65(6), 386–408. از متون بنیادی شبکههای عصبی یادگیرنده.
- Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323, 533–536. از مهمترین مقالات در احیای شبکههای عصبی چندلایه و backpropagation.
- Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet Classification with Deep Convolutional Neural Networks. Advances in Neural Information Processing Systems. مقاله AlexNet و یکی از نقاط عطف انقلاب یادگیری عمیق.
- Silver, D., et al. (2017). Mastering the Game of Go without Human Knowledge. Nature. پژوهش AlphaGo Zero درباره یادگیری تقویتی و self-play.
- Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention Is All You Need. مقاله معرفی معماری Transformer که بنیان بخش بزرگی از مدلهای زبانی و مولد امروزی شد.
- Brown, T. B., et al. (2020). Language Models are Few-Shot Learners. پژوهش GPT-3 درباره رابطه افزایش مقیاس مدلهای زبانی با قابلیتهای عمومیتر و few-shot learning.
- Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback. پژوهش InstructGPT و استفاده از بازخورد انسانی برای همسوسازی رفتار مدلهای زبان.
- OpenAI (2022). Introducing ChatGPT. اعلام انتشار عمومی ChatGPT در ۳۰ نوامبر ۲۰۲۲.
- Stanford Institute for Human-Centered Artificial Intelligence (2026). AI Index Report 2026. گزارش جامع درباره وضعیت فنی، اقتصادی، علمی و اجتماعی هوش مصنوعی تا سال ۲۰۲۶.
This article traces the history of artificial intelligence from its theoretical foundations through developments in 2026.
A History of Artificial Intelligence: From the Dream of a Thinking Machine to Generative AI, Reasoning Models, and Intelligent Agents
Artificial intelligence has entered everyday life, the economy, education, scientific research, industry, media, and politics so rapidly that it can seem like a phenomenon of only the last few years. The release of ChatGPT at the end of 2022, the spread of generative image and video models, the development of multimodal systems, and then the arrival of systems able to reason, write programs, use tools, and independently carry out parts of a workflow have unquestionably accelerated this transformation. Yet what we now call artificial intelligence is the result of more than eight decades of scientific development, with theoretical roots that precede modern electronic computers. The history of AI is the history of the convergence of mathematical logic, neuroscience, computation theory, statistics, computer engineering, linguistics, cognitive psychology, and, more recently, data science and the digital economy.
At the center of this history lies an old question: does what humans call “thinking” possess a structure that can be converted into computable operations? If so, can a machine do more than calculate numbers—can it recognize patterns, understand language, learn from experience, make decisions, and display behavior we would call intelligent?
The answers offered over the past eighty years have changed. Early researchers thought intelligence could be expressed as logical rules supplied to a machine. It later became clear that many dimensions of human intelligence cannot easily be written as explicit rules. Attention therefore shifted from programming the rules of intelligence to learning from data. Deep neural networks then made it possible to extract highly complex patterns. The Transformer architecture of 2017 opened the way to foundation models and very large language models. From the early 2020s, AI moved from being primarily a specialist tool toward becoming a general technology through which people could communicate in natural language. The current transition is moving it from “a machine that answers” toward “a system that can act to achieve a goal.”
To understand this path, we must begin before the term artificial intelligence existed. In the 1930s, British mathematician Alan Turing posed one of computer science’s most fundamental questions: what, in principle, is computable? His theoretical model, later called the Turing machine, gave computation a precise mathematical definition. Its importance for AI was that it established a theoretical framework for operations performed by a machine: if a process can be expressed as a definite sequence of operations, a general computing machine can in principle execute it.
In 1943, Warren McCulloch, a neuroscientist and psychiatrist, and Walter Pitts, a young logician, published “A Logical Calculus of the Ideas Immanent in Nervous Activity.” They proposed an abstract model of a neuron in which simple computational units could switch on or off according to their inputs and connect into networks. Their model differed greatly from present-day neural networks, but it was historically foundational because it established a formal relationship between neural activity and computational logic. It introduced the possibility that some functions of the nervous system might be modeled by networks of computational elements.
In 1950, Turing moved the question forward in “Computing Machinery and Intelligence,” published in Mind. Instead of becoming trapped in philosophical definitions of “machine” and “thought,” he proposed evaluating behavior through the “imitation game,” later known as the Turing test. In its simplest interpretation, if a human conversing through text cannot reliably determine whether the other participant is a human or a machine, the machine has displayed behavior that may practically be called intelligent. Turing’s contribution was not merely a test. He examined many objections to machine intelligence that remain familiar today and established the relationship between computation and intelligence as a serious research problem.
The term “artificial intelligence,” however, had not yet been coined. The event commonly treated as the formal birth of AI as a scientific field was the Dartmouth Summer Research Project of 1956. Its 1955 proposal was prepared by John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon. McCarthy used “artificial intelligence” to name the new field. Their conjecture was ambitious: every aspect of learning or intelligence could, in principle, be described precisely enough for a machine to simulate it. The 1956 gathering brought together researchers who became major founders of the field, making 1956 its conventional date of birth.
The 1950s and 1960s were years of extraordinary optimism. Early successes encouraged the belief that genuinely intelligent machines might be close. Allen Newell, Herbert Simon, and Cliff Shaw developed the Logic Theorist, which proved some theorems in mathematical logic. The General Problem Solver followed, attempting to provide general methods for solving problems. These systems embodied what became known as symbolic AI: the world could be represented through symbolic objects, concepts, and relations, while reasoning could be implemented as logical rules. Given sufficient knowledge and rules of inference, a machine could draw conclusions.
Symbolic AI worked well on constrained problems but faced a fundamental limitation. People perform many everyday activities without possessing an explicit rulebook. Recognizing a friend in the street, understanding a joke, detecting a speaker’s tone, or distinguishing thousands of objects cannot easily be converted into endless “if–then” statements. Much of human intelligence arises through experience and pattern learning that people themselves cannot always express as rules.
Alongside symbolic AI, another path was developing. In the late 1950s, Frank Rosenblatt created the Perceptron, a neuron-inspired system able to adjust its connection weights from training examples. Instead of requiring a programmer to determine every decision rule in advance, it could learn relationships from sample data. This became the crucial distinction between traditional programming and machine learning: rather than telling the computer exactly what to do for every input, one supplies examples from which it discovers a decision pattern. Rosenblatt’s 1958 paper in Psychological Review became a direct ancestor of modern neural networks.
Hardware and theoretical limitations soon became obvious. Computers had little processing power or memory, and digital data was scarce. Researchers also discovered that tasks easy for people could be extremely difficult for machines. A young child can learn to distinguish cats from dogs from only a few examples, whereas machines historically needed vast datasets. Translation, speech understanding, object recognition, and common-sense reasoning about real situations proved far harder than proving narrowly defined logical theorems.
During the 1970s, the gap between early promises and practical achievements produced cuts in funding and support. The period became known as an “AI winter.” AI has experienced more than one such cycle of excitement and disappointment. These cycles demonstrate that its progress has been neither linear nor inevitable; each advance has required a particular combination of scientific, hardware, and economic developments.
AI gained attention again in the 1980s through expert systems. These programs stored the knowledge of specialists as large collections of rules. A medical system could assess possible diagnoses from symptoms and test results, while an industrial system could recommend equipment configurations. MYCIN in medicine and XCON in computing became well-known examples. Expert systems showed that AI could create economic value in restricted domains, but maintaining thousands of rules became costly. Knowledge had to be extracted from experts, made explicit, recorded, and continually updated. The real world was too complex and dynamic to remain confined within fixed rules.
The same decade brought an important revival of neural networks. In 1986, David Rumelhart, Geoffrey Hinton, and Ronald Williams published “Learning Representations by Back-Propagating Errors” in Nature. Backpropagation measures the difference between a network’s output and the desired output, then carries the error backward through the layers so connection weights can be adjusted to reduce it. This method later became a principal foundation for training deep neural networks.
Neural networks nevertheless remained limited until three elements matured together: data, computing power, and algorithms. From the 1990s, and especially during the 2000s, the internet, smartphones, e-commerce, social media, digital systems, and sensors generated enormous datasets. Processor performance multiplied, and graphics processing units—originally designed for computer graphics—proved highly effective for the parallel operations required by neural networks. Algorithms, network-training methods, large datasets, and software tools also improved.
The meaning of AI gradually changed. Rather than having researchers directly supply all knowledge, machines extracted statistical relationships from data. In a traditional spam filter, a programmer might write rules about particular words or senders. In machine learning, thousands or millions of emails labeled “spam” or “normal” are supplied, and the model learns the relevant combination of features. Machine knowledge thus shifted from something directly written by a programmer to something extracted from data.
In 1997, IBM’s Deep Blue defeated world chess champion Garry Kasparov in a formal match. Chess had long symbolized strategic thought and human intelligence, so the victory had enormous public significance. Yet Deep Blue was fundamentally unlike today’s general models. It was designed specifically for chess and relied on extensive search, position evaluation, and specialized knowledge. It showed that a machine could surpass humans in one bounded domain, not that it possessed flexible general intelligence.
The next great turning point arrived in 2012. Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton trained a deep neural network for the ImageNet image-recognition competition. AlexNet, trained on more than a million images, dramatically reduced error compared with earlier methods. It demonstrated that deep neural networks, given enough data and computation, could learn complex visual features for themselves. This marked the practical beginning of the deep-learning explosion of the 2010s, after which deep networks rapidly became dominant in vision, speech, translation, language processing, medicine, and many other fields.
Deep learning uses neural networks with many layers, each producing a more complex representation of data. In image recognition, early layers may identify edges; intermediate layers learn shapes or textures; deeper layers recognize structures such as eyes, faces, animals, or vehicles. The engineer no longer needs to define all relevant features manually—the network extracts them from data.
Four years after AlexNet, another symbolic event changed public perceptions. In March 2016, DeepMind’s AlphaGo defeated Lee Sedol, one of the world’s leading Go players, four games to one. Because Go has an astronomical number of possible states, it had long been expected to resist machines much longer than chess. AlphaGo combined neural networks, reinforcement learning, and tree search, showing that learning and computation could reach solutions beyond brute-force enumeration.
AlphaGo Zero and AlphaZero went further. Through reinforcement learning and self-play, systems learned powerful strategies without relying on vast archives of human games. Starting with only the rules, AlphaGo Zero improved its network by playing against itself and effectively became its own teacher. Machine learning had entered a new phase: the machine could learn not only from human-generated data but also from experience it generated itself.
The development that led most directly to today’s large language models came in 2017, when Google researchers published “Attention Is All You Need” and introduced the Transformer. Earlier language systems often relied on recurrent architectures that struggled with long sequences. The Transformer used attention, allowing a model to assess how every part of a text relates to other parts while enabling far greater parallelism during training. Initially demonstrated in machine translation, the architecture later became the foundation of most large language models, multimodal models, and contemporary generative systems.
After the Transformer, the idea of a foundation model grew in importance. Instead of building separate small models for translation, summarization, and question answering, one very large model could be trained on massive bodies of information and then used for many tasks. GPT-2 in 2019 showed that a large language model trained on broad text could produce coherent passages and display abilities in answering questions, summarizing, and translating without separate task-specific training.
GPT-3, introduced in 2020 with 175 billion parameters, created another leap. “Language Models Are Few-Shot Learners” showed that greater scale could expand a model’s ability to perform many tasks from a few examples or natural-language instructions. This changed the human–machine interface. In conventional software, changing behavior requires changing code; with large language models, users could increasingly guide behavior through ordinary language. The prompt became a new interface between human intention and computation.
Another technical and social development concerned training models to respond helpfully to people. Predicting the next token alone does not guarantee that a system will understand a request or provide a useful response. Reinforcement learning from human feedback (RLHF) allowed models to learn from human preferences about answer quality. InstructGPT became an important example; research published in 2022 showed that human-feedback tuning could improve instruction-following and interaction.
ChatGPT was released on November 30, 2022. Most of its scientific components already existed, but socially it was a turning point. Its conversational interface concealed the complexity of large models. A person with no programming knowledge could use everyday language to request a summary, a program, an explanation, an edit, ideas, or data analysis. AI entered what might be called the democratization of the AI interface. Just as graphical interfaces made personal computing broadly accessible, conversational models made ordinary language a principal way to direct digital systems.
Models then became increasingly multimodal. Humans do not experience the world through text alone: image, sound, language, movement, and physical surroundings all participate in cognition. Multimodal systems can receive or produce information in several forms. They can analyze an image and discuss it, explain a chart, hear and answer speech, produce images, or combine textual and visual information. The boundaries among image recognition, language models, audio systems, and software have consequently begun to blur.
The next major direction was the rise of reasoning models. Early language models produced fluent text but often failed on complex, multistep problems in mathematics or programming. Newer AI systems place greater emphasis on problem solving: decomposing a goal into steps, considering alternatives, using external tools, executing calculations, and evaluating results. Stanford’s 2026 AI Index reports rapid gains in difficult scientific, mathematical, multimodal-reasoning, and coding evaluations. Some frontier models have reached or surpassed human benchmarks on selected tests, while performance on certain software-engineering benchmarks has risen sharply within a short period.
Alongside reasoning models, agentic AI has become a major development path. The difference between a chatbot and an agent is fundamental. A chatbot normally receives a question and returns an answer. An intelligent agent can receive a goal, break it into stages, decide on an order of work, use search engines, databases, applications, files, or code, inspect the result of one stage, and choose the next. AI is changing from an answering machine into a task-performing system. Yet agentic systems still face important limits in reliability, oversight, and autonomy.
This transition may prove even more economically consequential than chatbots. When AI only produces text, it is a productivity tool. When it can conduct part of a production, research, programming, administrative, support, or analytical process from beginning to end, it becomes an active component in organizing work. The future question is therefore not only what a machine knows, but how much of the cycle from goal to decision and from decision to action may be entrusted to intelligent systems.
By 2026, describing AI as merely another branch of computer science is no longer adequate. It is becoming a general infrastructure for cognitive activity. Advanced models can process text, images, audio, and code; assist scientific research; analyze large datasets; participate in software development; solve mathematical problems; connect to external tools; and pursue multistep tasks as agents. Stanford’s 2026 AI Index also reports that industry produced more than 90 percent of notable frontier models in 2025. The center of gravity in frontier AI has shifted from universities and public laboratories toward companies able to finance enormous requirements for chips, data centers, electricity, networks, and specialized labor.
Here the technical history of AI joins its economic history. Contemporary AI is not simply an algorithm; it is the combination of algorithms, data, chips, energy, data centers, communications networks, and immense capital. Discussion of AI therefore necessarily includes ownership of its infrastructure. Advanced models may be trained on humanity’s accumulated knowledge—books, research, software, images, languages, and cultural works—while the capacity to build and control frontier systems remains concentrated in relatively few large institutions. A historical contradiction emerges: AI’s intellectual raw material is profoundly social, while the means of processing it into economic power may be highly concentrated.
The question of the future is not only what machines can do. We must ask who owns the machines, data, models, chips, and infrastructure; who determines system objectives; who has access; who receives the benefits of increased productivity; and who bears the social costs.
From this perspective, AI history is also part of the history of productive forces. The steam engine multiplied human muscular power. Electric motors and assembly lines enabled mass production. Computers automated calculation and information processing. The internet enabled near-instant transmission and sharing of information at global scale. Artificial intelligence is now entering another realm: part of cognitive labor.
Cognitive labor includes writing, translation, analysis, design, programming, pattern recognition, planning, research, and some decision support—activities until recently regarded as exclusively or almost exclusively human. AI will not necessarily replace all of them, but it changes the relationship between people and their tools of production. Industrial machines did not instantly eliminate physical labor, but transformed its character and organization; AI is likewise changing the structure of intellectual work.
One important mistake must be avoided: intelligence and consciousness are not the same. The remarkable performance of contemporary models shows that machines can produce outputs once treated as signs of human intelligence, but it does not by itself show that machines possess consciousness or subjective experience. A model may write thousands of pages about pain, love, fear, or death without evidence that it feels any of them. It may reason about a concept without possessing self-aware experience comparable to that of a human being. We must distinguish information processing, learning, intelligence, cognition, and consciousness.
This distinction matters socially. A society may possess an unprecedented level of computational intelligence without becoming more conscious. A powerful system can support drug discovery, education, economic planning, the reduction of waste, and wider access to knowledge. The same or similar technology can be used for mass surveillance, information manipulation, warfare, monopoly, or social control. Technology does not determine its own ethical and social direction. Human beings, institutions, and social structures determine how technical capacity is used.
AI history can therefore be read at two levels. At the technical level, mathematical logic led to machine computation; computation to symbolic AI; symbolic AI to machine learning; machine learning to deep networks; deep networks to the Transformer; the Transformer to foundation models; and foundation models to generative and multimodal AI, reasoning systems, and intelligent agents.
At a deeper level, humans first gave machines rules; then supplied data from which patterns could be learned; machines created increasingly complex representations of information; large models absorbed part of society’s accumulated knowledge into statistical structures; and agentic systems now attempt to turn that knowledge into action.
Rule → Data → Information → Pattern → Model → Operational Knowledge → Reasoning → Action
This chain explains why contemporary AI differs from earlier technology. A tool is no longer used only to execute a fixed instruction. It increasingly participates in interpreting the objective, finding a path, and selecting some actions.
The question Turing asked in 1950—can machines think?—has not disappeared, but by 2026 it is no longer the only important question. Machines can already produce behavior treated as intelligent across many domains. The larger social question is what humanity will do with this new capacity.
Will AI become primarily an instrument for concentrating wealth and power, or can it expand knowledge and social participation? Will its productivity gains reduce working time and improve life, or eliminate jobs without distributing the benefits? Will social knowledge and data become public resources for human empowerment or raw material for new monopolies? Will AI make power more transparent, or citizens more observable? Most importantly, will greater machine “intelligence” contribute to greater human and social “consciousness”?
This is where the history of AI connects to the broader discussion of A New Social Order in the Age of Consciousness. The Industrial Revolution placed ownership of factories and productive tools at the center of social conflict. The digital revolution foregrounded ownership of information and networks. The AI revolution now raises the issue of owning and controlling infrastructures that process not only material goods, but knowledge, decisions, and cognitive activity.
AI cannot be understood solely as a technical invention. It is becoming one of the decisive productive forces of the twenty-first century. As earlier technologies transformed economy and society, AI will likely reshape work, education, ownership, power, politics, and even our definition of human skill.
The most important chapter in AI history may therefore remain unwritten. Its first chapter asked whether machines could calculate. The next asked whether they could reason. Later came learning, vision, hearing, and language. Today the question is moving toward problem solving and relatively autonomous action. Yet the next chapter will be less a question about machines than a question about society.
Until now, scientists have principally asked: How intelligent can a machine become?
The decisive question of the coming age may be: What will human society do with this intelligence?
The answer to the first question has largely been written in laboratories, universities, and data centers. Engineers alone cannot determine the answer to the second. It will depend on economics, politics, ethics, law, culture, ownership, social participation, and ultimately the level of society’s consciousness.
In this sense, the history of artificial intelligence is no longer only the history of intelligent machines. It is becoming part of the history of the transformation of human society itself.
References and Further Reading
- Turing, A. M. (1950). Computing Machinery and Intelligence. Mind, 59(236), 433–460.
- McCulloch, W. S., & Pitts, W. (1943). A Logical Calculus of the Ideas Immanent in Nervous Activity. Bulletin of Mathematical Biophysics, 5, 115–133.
- McCarthy, J., Minsky, M., Rochester, N., & Shannon, C. (1955). A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence.
- Rosenblatt, F. (1958). The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain. Psychological Review, 65(6), 386–408.
- Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning Representations by Back-Propagating Errors. Nature, 323, 533–536.
- Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet Classification with Deep Convolutional Neural Networks. Advances in Neural Information Processing Systems.
- Silver, D., et al. (2017). Mastering the Game of Go without Human Knowledge. Nature.
- Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention Is All You Need.
- Brown, T. B., et al. (2020). Language Models Are Few-Shot Learners.
- Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback.
- OpenAI (2022). Introducing ChatGPT.
- Stanford Institute for Human-Centered Artificial Intelligence (2026). AI Index Report 2026.