Age of Consciousness عصر آگاهی

تاریخ هوش مصنوعی؛ از رؤیای ماشین متفکر تا ظهور هوش مولد | A History of Artificial Intelligence

این مقاله روند تاریخی هوش مصنوعی را از بنیان‌های نظری تا تحولات سال ۲۰۲۶ بررسی می‌کند.

تاریخ هوش مصنوعی؛ از رؤیای ماشین متفکر تا ظهور هوش مولد، مدل‌های استدلالی و عامل‌های هوشمند

هوش مصنوعی امروز چنان سریع وارد زندگی روزمره، اقتصاد، آموزش، پژوهش علمی، صنعت، رسانه و سیاست شده است که ممکن است این تصور ایجاد شود که با پدیده‌ای متعلق به همین چند سال اخیر روبه‌رو هستیم. ظهور ChatGPT در پایان سال ۲۰۲۲، گسترش مدل‌های مولد تصویر و ویدئو، توسعه مدل‌های چندوجهی و سپس ظهور سیستم‌هایی که می‌توانند استدلال کنند، برنامه بنویسند، از ابزارهای مختلف استفاده کنند و بخشی از یک فرایند کاری را مستقلاً پیش ببرند، بدون تردید شتابی بی‌سابقه به این تحول داده است. با این همه، آنچه امروز «هوش مصنوعی» می‌نامیم حاصل بیش از هشت دهه تحول علمی است و ریشه‌های نظری آن حتی به سال‌های پیش از پیدایش رایانه‌های الکترونیکی مدرن بازمی‌گردد. تاریخ هوش مصنوعی در واقع تاریخ تلاقی چند رشته است: منطق ریاضی، علوم اعصاب، نظریه محاسبه، آمار، مهندسی کامپیوتر، زبان‌شناسی، روان‌شناسی شناختی و، در دهه‌های اخیر، علوم داده و اقتصاد دیجیتال.

در مرکز این تاریخ یک پرسش بسیار قدیمی قرار دارد: آیا آنچه انسان «اندیشیدن» می‌نامد، دارای ساختاری است که بتوان آن را به مجموعه‌ای از عملیات قابل محاسبه تبدیل کرد؟ اگر پاسخ مثبت باشد، آیا یک ماشین می‌تواند نه فقط اعداد را محاسبه کند، بلکه الگوها را تشخیص دهد، زبان را بفهمد، از تجربه بیاموزد، تصمیم بگیرد و حتی رفتارهایی از خود نشان دهد که ما آنها را هوشمندانه می‌نامیم؟

پاسخ‌هایی که طی هشتاد سال گذشته به این پرسش داده شده‌اند یکسان نبوده‌اند. در نخستین دوره تصور می‌شد هوش را می‌توان به قواعد منطقی تبدیل کرد و آن قواعد را به ماشین داد. سپس معلوم شد که بسیاری از جنبه‌های هوش انسانی را نمی‌توان به‌آسانی در قالب قواعد صریح نوشت. در نتیجه، مرکز توجه از «برنامه‌ریزی قواعد هوش» به «یادگیری از داده» منتقل شد. بعد شبکه‌های عصبی عمیق امکان استخراج الگوهای بسیار پیچیده را فراهم کردند. ظهور Transformer در سال ۲۰۱۷ راه را برای ساخت مدل‌های بنیادین و مدل‌های زبانی بسیار بزرگ گشود. سرانجام، از اوایل دهه ۲۰۲۰، هوش مصنوعی از مرحله‌ای که عمدتاً یک ابزار تخصصی بود به فناوری‌ای عمومی تبدیل شد که قادر است با انسان از طریق زبان طبیعی ارتباط برقرار کند. تحول کنونی نیز آن را از «ماشینی که پاسخ می‌دهد» به سوی «سیستمی که می‌تواند برای رسیدن به یک هدف اقدام کند» سوق می‌دهد.

برای فهم این مسیر باید از زمانی آغاز کرد که هنوز اصطلاح «هوش مصنوعی» وجود نداشت. در دهه ۱۹۳۰، آلن تورینگ، ریاضیدان بریتانیایی، یکی از بنیادی‌ترین پرسش‌های علوم کامپیوتر را مطرح کرد: اصولاً چه چیزی قابل محاسبه است؟ مدل نظری او که بعدها «ماشین تورینگ» نام گرفت، نشان داد که می‌توان مفهوم محاسبه را به شکلی دقیق و ریاضی تعریف کرد. اهمیت این دستاورد برای تاریخ هوش مصنوعی در این بود که مرزی نظری برای انجام عملیات توسط ماشین فراهم کرد. اگر فرایندی را بتوان به توالی مشخصی از عملیات تبدیل کرد، در اصل یک ماشین عمومی محاسبه نیز می‌تواند آن را اجرا کند.

گام مهم بعدی در سال ۱۹۴۳ برداشته شد؛ زمانی که Warren McCulloch، عصب‌شناس و روان‌پزشک، و Walter Pitts، منطق‌دان جوان، مقاله مشهور خود با عنوان «A Logical Calculus of the Ideas Immanent in Nervous Activity» را منتشر کردند. آنان مدلی انتزاعی از نورون ارائه کردند که در آن واحدهای ساده محاسباتی می‌توانستند بسته به ورودی‌هایشان فعال یا غیرفعال شوند و در قالب شبکه‌هایی به یکدیگر متصل گردند. این مدل با شبکه‌های عصبی امروزی تفاوت‌های زیادی داشت، اما از نظر تاریخی بنیادی بود، زیرا برای نخستین بار ارتباطی رسمی میان عملکرد شبکه عصبی و منطق محاسباتی برقرار می‌کرد. به بیان دیگر، این تصور شکل گرفت که شاید بتوان برخی عملکردهای سیستم عصبی را با شبکه‌ای از عناصر محاسباتی مدل‌سازی کرد. مقاله McCulloch و Pitts که در Bulletin of Mathematical Biophysics منتشر شد، بعدها یکی از پایه‌های نظری شبکه‌های عصبی مصنوعی به شمار آمد.

در سال ۱۹۵۰ تورینگ پرسش را یک مرحله جلوتر برد. مقاله او با عنوان «Computing Machinery and Intelligence» در مجله Mind با این پرسش آغاز شد که آیا ماشین‌ها می‌توانند فکر کنند. تورینگ به جای گرفتار شدن در تعریف فلسفی واژه‌های «ماشین» و «تفکر»، پیشنهاد کرد مسئله را از طریق رفتار بررسی کنیم. او «بازی تقلید» را مطرح کرد؛ آزمایشی که بعدها به آزمون تورینگ شهرت یافت. در ساده‌ترین تعبیر آن، اگر یک انسان از طریق گفت‌وگوی متنی نتواند با اطمینان تشخیص دهد که طرف مقابل انسان است یا ماشین، ماشین رفتاری از خود نشان داده که از نظر عملی می‌توان آن را هوشمندانه تلقی کرد. اهمیت مقاله تورینگ فقط در ارائه یک آزمون نبود؛ او بسیاری از استدلال‌هایی را که هنوز هم درباره امکان هوش ماشینی مطرح می‌شوند بررسی کرد و نشان داد که مسئله رابطه میان محاسبه و هوش باید به یک مسئله پژوهشی جدی تبدیل شود.

با این همه، اصطلاح «Artificial Intelligence» هنوز وجود نداشت. نقطه‌ای که معمولاً تولد رسمی هوش مصنوعی به‌عنوان یک رشته علمی تلقی می‌شود، پروژه تابستانی Dartmouth در سال ۱۹۵۶ است. پیشنهاد اولیه این پروژه در سال ۱۹۵۵ توسط John McCarthy، Marvin Minsky، Nathaniel Rochester و Claude Shannon تهیه شده بود. McCarthy اصطلاح «Artificial Intelligence» را برای نامیدن حوزه جدید به کار برد. فرض اساسی پیشنهاد Dartmouth بسیار بلندپروازانه بود: این پژوهشگران تصور می‌کردند هر جنبه‌ای از یادگیری یا دیگر ویژگی‌های هوش را می‌توان در اصل آن‌قدر دقیق توصیف کرد که ماشین قادر به شبیه‌سازی آن باشد. گردهمایی تابستان ۱۹۵۶ در Dartmouth گروهی از افرادی را کنار هم آورد که بعدها به بنیان‌گذاران اصلی رشته AI تبدیل شدند و از همین رو سال ۱۹۵۶ معمولاً تاریخ تولد رسمی این رشته محسوب می‌شود.

دهه‌های ۱۹۵۰ و ۱۹۶۰ دوران خوش‌بینی فوق‌العاده بود. رایانه‌ها تازه در حال شکل‌گیری بودند و نخستین موفقیت‌ها این تصور را تقویت می‌کردند که شاید ایجاد ماشین‌های واقعاً هوشمند چندان دور نباشد. یکی از نخستین برنامه‌های مشهور، Logic Theorist بود که Allen Newell، Herbert Simon و Cliff Shaw آن را توسعه دادند. این برنامه می‌توانست برخی قضایای منطق ریاضی را اثبات کند. کمی بعد General Problem Solver ساخته شد و تلاش کرد روش‌هایی عمومی برای حل مسائل ارائه دهد. در این دوران رویکردی شکل گرفت که بعدها «هوش مصنوعی نمادین» یا Symbolic AI نامیده شد. فرض آن این بود که جهان را می‌توان در قالب اشیا، مفاهیم و روابط نمادین توصیف کرد و استدلال را نیز به قواعد منطقی تبدیل نمود. اگر دانش کافی درباره یک حوزه به ماشین داده شود و قواعد استنتاج نیز مشخص باشند، ماشین قادر خواهد بود نتیجه‌گیری کند.

این رویکرد برای مسائل محدود بسیار موفق بود، اما محدودیتی اساسی داشت: انسان‌ها بسیاری از کارهای روزمره خود را بدون داشتن مجموعه‌ای صریح از قواعد انجام می‌دهند. برای نمونه، تشخیص چهره یک دوست در خیابان، فهم یک شوخی، تشخیص لحن گوینده یا تمایز دادن میان هزاران شیء متفاوت به‌سادگی قابل تبدیل به هزاران جمله «اگر… آنگاه…» نیست. هوش انسانی بخش عظیمی از توان خود را از تجربه و یادگیری الگوهایی به دست می‌آورد که انسان حتی نمی‌تواند همیشه آنها را به‌صورت قواعد صریح بیان کند.

هم‌زمان با رویکرد نمادین، مسیر دیگری نیز در حال شکل‌گیری بود که بعدها نقشی تعیین‌کننده یافت. Frank Rosenblatt در اواخر دهه ۱۹۵۰ مدل Perceptron را توسعه داد؛ سامانه‌ای الهام‌گرفته از نورون که می‌توانست بر اساس نمونه‌های آموزشی وزن‌های ارتباطی خود را تغییر دهد. اهمیت پرسپترون از این جهت بود که به جای آنکه برنامه‌نویس تمام قواعد تصمیم‌گیری را از پیش تعیین کند، سیستم می‌توانست برخی روابط را از داده‌های نمونه یاد بگیرد. این همان تفاوت بنیادی است که بعدها میان برنامه‌نویسی سنتی و یادگیری ماشین شکل گرفت: به جای اینکه انسان دقیقاً به رایانه بگوید برای هر ورودی چه کند، نمونه‌هایی در اختیار آن گذاشته می‌شود تا خود الگویی برای تصمیم‌گیری پیدا کند. پژوهش Rosenblatt در سال ۱۹۵۸ در Psychological Review منتشر شد و یکی از اجداد مستقیم شبکه‌های عصبی امروزی به شمار می‌رود.

اما محدودیت‌های سخت‌افزاری و نظری خیلی زود خود را نشان دادند. رایانه‌ها قدرت پردازش اندکی داشتند، حافظه بسیار محدود بود و حجم داده دیجیتال با امروز قابل مقایسه نبود. مهم‌تر از همه، پژوهشگران دریافتند مسائلی که برای انسان بسیار ساده به نظر می‌رسند می‌توانند برای ماشین فوق‌العاده دشوار باشند. انسان کودک خردسالی را با چند نمونه قادر می‌کند گربه را از سگ تشخیص دهد، در حالی که آموزش ماشین برای همان کار به حجم عظیمی از داده نیاز داشت. ترجمه زبان، درک گفتار، تشخیص اشیا و استدلال درباره موقعیت‌های جهان واقعی بسیار دشوارتر از اثبات قضایای محدود منطقی بودند.

در دهه ۱۹۷۰ فاصله میان وعده‌های اولیه و دستاوردهای عملی باعث کاهش حمایت مالی در برخی مراکز شد. این دوره بعدها «زمستان هوش مصنوعی» نام گرفت. اصطلاح AI Winter به دوره‌هایی اشاره دارد که پس از موج‌های خوش‌بینی و سرمایه‌گذاری، عدم تحقق انتظارات موجب کاهش بودجه و توجه عمومی شد. هوش مصنوعی در تاریخ خود بیش از یک چنین دوره‌ای را تجربه کرد و همین فراز و فرودها نشان می‌دهد که پیشرفت AI نه خطی بوده و نه اجتناب‌ناپذیر؛ هر جهش آن به مجموعه‌ای از پیشرفت‌های علمی، سخت‌افزاری و اقتصادی نیاز داشته است.

در دهه ۱۹۸۰ هوش مصنوعی بار دیگر، این بار با «سیستم‌های خبره» یا Expert Systems، مورد توجه قرار گرفت. در این سیستم‌ها دانش متخصصان یک حوزه در قالب مجموعه بزرگی از قواعد ذخیره می‌شد. برای مثال یک سیستم پزشکی می‌توانست بر اساس ترکیبی از علائم و نتایج آزمایش احتمالاتی درباره بیماری ارائه دهد یا یک سیستم صنعتی برای پیکربندی تجهیزات پیشنهادهایی مطرح کند. MYCIN در پزشکی و XCON در صنعت کامپیوتر نمونه‌های شناخته‌شده این رویکرد بودند. سیستم‌های خبره نشان دادند AI می‌تواند در حوزه‌های محدود اقتصادی ارزش عملی داشته باشد، اما مشکل نگهداری هزاران قاعده به تدریج آشکار شد. «مهندسی دانش» به فرایندی پرهزینه تبدیل می‌شد: دانش متخصص باید استخراج، صریح، ثبت و دائماً به‌روز می‌شد. جهان واقعی بیش از اندازه پیچیده و پویا بود که همیشه بتوان آن را در مجموعه‌ای از قواعد ثابت محصور کرد.

یکی از تحولات علمی مهم همین دهه، احیای شبکه‌های عصبی بود. در سال ۱۹۸۶ David Rumelhart، Geoffrey Hinton و Ronald Williams مقاله مشهور «Learning representations by back-propagating errors» را در Nature منتشر کردند. این مقاله استفاده مؤثر از الگوریتم backpropagation را برای تنظیم وزن‌های شبکه‌های چندلایه توضیح می‌داد. اصل کار آن است که شبکه ابتدا خروجی تولید می‌کند، اختلاف میان خروجی واقعی و خروجی مورد انتظار اندازه‌گیری می‌شود و سپس خطا از لایه‌های پایانی به عقب منتقل می‌گردد تا وزن ارتباطات به نحوی تغییر کنند که خطا کاهش یابد. این روش بعدها به یکی از ستون‌های اصلی آموزش شبکه‌های عصبی عمیق تبدیل شد.

با این همه، تا مدت‌ها شبکه‌های عصبی محدود باقی ماندند. برای آنکه توان واقعی آنها آشکار شود سه عامل باید هم‌زمان رشد می‌کردند: داده، قدرت محاسباتی و الگوریتم. این سه عامل از دهه ۱۹۹۰ و به‌ویژه در دهه ۲۰۰۰ به تدریج به یکدیگر رسیدند. اینترنت، تلفن‌های هوشمند، تجارت الکترونیک، شبکه‌های اجتماعی، سیستم‌های دیجیتال و حسگرها حجم عظیمی از داده ایجاد کردند. توان پردازنده‌ها چندین مرتبه افزایش یافت و GPUها، که در اصل برای پردازش گرافیک توسعه یافته بودند، مشخص شد برای عملیات موازی مورد نیاز شبکه‌های عصبی بسیار مناسب‌اند. هم‌زمان الگوریتم‌ها، روش‌های تنظیم شبکه، مجموعه‌داده‌های بزرگ و ابزارهای نرم‌افزاری بهبود یافتند.

در این دوره معنای هوش مصنوعی نیز آرام‌آرام تغییر کرد. به جای آنکه پژوهشگر همه دانش را به ماشین بدهد، ماشین با استفاده از داده‌های بزرگ به استخراج روابط آماری می‌پرداخت. این همان حوزه‌ای است که Machine Learning یا یادگیری ماشین نام گرفت. برای مثال، در روش سنتی تشخیص ایمیل هرزنامه، برنامه‌نویس ممکن بود ده‌ها قاعده تعریف کند: اگر این کلمات وجود داشتند یا فرستنده چنین ویژگی‌هایی داشت احتمال Spam بیشتر است. در یادگیری ماشین، هزاران یا میلیون‌ها نمونه ایمیل با برچسب «هرزنامه» یا «عادی» در اختیار الگوریتم قرار می‌گیرد و مدل خودش ترکیبی از ویژگی‌های مؤثر را یاد می‌گیرد. بنابراین تحول مهمی رخ داد: دانش ماشین به‌تدریج از چیزی که مستقیماً توسط برنامه‌نویس نوشته می‌شد، به چیزی تبدیل شد که از داده استخراج می‌شد.

یکی از نمادهای دوره میانی AI در سال ۱۹۹۷ رقم خورد، زمانی که سیستم Deep Blue شرکت IBM توانست Garry Kasparov، قهرمان جهان شطرنج، را در یک مسابقه رسمی شکست دهد. در افکار عمومی این رویداد بسیار مهم بود، زیرا شطرنج قرن‌ها نمادی از تفکر استراتژیک و هوش انسانی محسوب می‌شد. با این حال Deep Blue با مدل‌های AI امروز تفاوت بنیادی داشت. این سیستم برای حوزه خاص شطرنج ساخته شده بود و از جست‌وجوی گسترده، ارزیابی موقعیت‌ها و دانش تخصصی استفاده می‌کرد. شکست Kasparov نشان داد ماشین می‌تواند در یک حوزه محدود از انسان پیشی بگیرد، اما به معنای وجود هوشی عمومی و انعطاف‌پذیر نبود.

نقطه عطف بزرگ بعدی در سال ۲۰۱۲ پدید آمد. Alex Krizhevsky، Ilya Sutskever و Geoffrey Hinton شبکه‌ای عمیق را برای مسابقه تشخیص تصویر ImageNet آموزش دادند. شبکه‌ای که بعدها AlexNet نامیده شد، روی بیش از یک میلیون تصویر آموزش دید و توانست میزان خطا را به شکل چشمگیری نسبت به روش‌های پیشین کاهش دهد. مقاله آنان در کنفرانس NeurIPS نشان داد شبکه‌های عصبی عمیق، در صورتی که داده و قدرت محاسباتی کافی در اختیار داشته باشند، قادرند ویژگی‌های پیچیده تصویر را خودشان یاد بگیرند. این رویداد در عمل آغاز انفجار Deep Learning در دهه ۲۰۱۰ بود. پس از آن شبکه‌های عمیق به سرعت در تشخیص تصویر، گفتار، ترجمه، پردازش زبان، پزشکی و بسیاری حوزه‌های دیگر به روش غالب تبدیل شدند.

Deep Learning یا یادگیری عمیق اساساً شبکه‌های عصبی چندلایه‌ای است که هر لایه قادر است نمایش پیچیده‌تری از داده ایجاد کند. برای نمونه، در یک سیستم تشخیص تصویر ممکن است لایه‌های اولیه الگوهای ساده‌ای مانند لبه‌ها را تشخیص دهند، لایه‌های میانی ترکیب‌هایی مانند اشکال یا بافت‌ها را بیاموزند و لایه‌های عمیق‌تر ساختارهایی مانند چشم، صورت، حیوان یا خودرو را تشخیص دهند. ویژگی مهم این تحول آن بود که مهندس دیگر لازم نبود تمام ویژگی‌های لازم برای تشخیص تصویر را دستی تعریف کند؛ شبکه آنها را از داده استخراج می‌کرد.

چهار سال پس از AlexNet، رویداد نمادین دیگری تصور عمومی از AI را تغییر داد. در مارس ۲۰۱۶ AlphaGo، سیستم شرکت DeepMind، در یک مسابقه پنج‌بازی Lee Sedol، یکی از برجسته‌ترین بازیکنان جهان در بازی Go را با نتیجه چهار بر یک شکست داد. Go به دلیل تعداد نجومی حالت‌های ممکن بسیار پیچیده‌تر از شطرنج است و مدت‌ها تصور می‌شد ماشین برای رسیدن به سطح قهرمانان انسانی در آن به زمان بسیار بیشتری نیاز دارد. AlphaGo از شبکه‌های عصبی، یادگیری تقویتی و جست‌وجوی درختی استفاده می‌کرد. پیروزی آن نشان داد ترکیب یادگیری و محاسبات می‌تواند به راه‌حل‌هایی برسد که به مراتب فراتر از جست‌وجوی خام تمام حالت‌های ممکن است.

نسخه‌های بعدی این ایده حتی فراتر رفتند. AlphaGo Zero و AlphaZero نشان دادند که سیستم می‌تواند با بازی کردن علیه خودش و از طریق Reinforcement Learning، بدون نیاز به انبوه بازی‌های انسانی، استراتژی‌های بسیار قدرتمندی بیاموزد. مقاله AlphaGo Zero در Nature توضیح داد که سیستم می‌توانست تنها بر اساس قواعد بازی و خودبازی، شبکه خود را بهبود دهد و عملاً «معلم خودش» شود. اینجا مفهوم یادگیری ماشین به مرحله تازه‌ای رسید: ماشین نه فقط از داده‌های تولیدشده توسط انسان، بلکه از تجربه‌ای که خودش ایجاد می‌کرد نیز می‌آموخت.

با این همه، انقلابی که مستقیماً به مدل‌های زبانی بزرگ امروز انجامید، در سال ۲۰۱۷ رخ داد. گروهی از پژوهشگران Google مقاله‌ای با عنوان «Attention Is All You Need» منتشر کردند و معماری Transformer را معرفی نمودند. تا آن زمان بسیاری از سیستم‌های پردازش زبان از شبکه‌های بازگشتی یا روش‌هایی استفاده می‌کردند که پردازش توالی‌های طولانی را دشوار می‌ساخت. Transformer بر سازوکاری به نام Attention تکیه کرد که به مدل اجازه می‌دهد هنگام پردازش هر بخش از متن، میزان ارتباط آن را با بخش‌های دیگر ارزیابی کند. مزیت مهم دیگر این بود که Transformer امکان موازی‌سازی بسیار بیشتری در جریان آموزش فراهم می‌کرد. نویسندگان مقاله نشان دادند این معماری در ترجمه ماشینی عملکردی بهتر و آموزش سریع‌تری نسبت به بسیاری از روش‌های قبلی دارد. اهمیت تاریخی این مقاله بعدها آشکار شد: معماری Transformer بنیان اصلی بسیاری از مدل‌های زبان بزرگ، مدل‌های چندوجهی و سیستم‌های مولد امروزی شد.

پس از Transformer مفهوم «مدل بنیادین» یا Foundation Model اهمیت فزاینده‌ای یافت. به جای ساخت یک مدل کوچک جداگانه برای ترجمه، یکی دیگر برای خلاصه‌سازی و مدل دیگری برای پاسخ به سؤال، می‌توان یک مدل بسیار بزرگ را روی حجم عظیمی از اطلاعات آموزش داد و سپس از همان مدل برای وظایف متعدد استفاده کرد. مدل‌های GPT یکی از مهم‌ترین مسیرهای این تحول بودند. GPT-2 در سال ۲۰۱۹ نشان داد که مدل زبانی بزرگ آموزش‌دیده روی متون وسیع می‌تواند متن‌های نسبتاً منسجم تولید کند و بدون آموزش جداگانه در وظایفی مانند پاسخ به سؤال، خلاصه‌سازی و ترجمه قابلیت‌هایی از خود نشان دهد.

در سال ۲۰۲۰ GPT-3 با ۱۷۵ میلیارد پارامتر جهش دیگری ایجاد کرد. مقاله «Language Models are Few-Shot Learners» نشان داد که صرف افزایش مقیاس مدل می‌تواند توانایی آن را برای انجام وظایف متعدد از طریق نمونه‌های اندک یا حتی دستورهای متنی افزایش دهد. این تحول از لحاظ رابط انسان و ماشین بسیار مهم بود. در برنامه‌نویسی سنتی، برای تغییر رفتار نرم‌افزار باید کد تغییر کند؛ اما در مدل‌های زبانی بزرگ، کاربر به تدریج می‌توانست تنها از طریق زبان طبیعی رفتار سیستم را هدایت کند. Prompt به نوعی رابط جدید میان خواست انسان و فرایند محاسباتی تبدیل شد.

در اینجا باید به یک تحول فنی و اجتماعی دیگر نیز توجه کرد: آموزش مدل برای «پاسخ مطلوب به انسان». پیش‌بینی توکن بعدی به خودی خود تضمین نمی‌کند که مدل سؤال کاربر را درست بفهمد یا پاسخ مفیدی بدهد. توسعه روش‌هایی مانند Reinforcement Learning from Human Feedback یا RLHF به مدل اجازه داد از ترجیحات انسانی درباره کیفیت پاسخ‌ها نیز یاد بگیرد. InstructGPT نمونه مهم این مسیر بود و تحقیقات منتشرشده در سال ۲۰۲۲ نشان دادند که تنظیم مدل با بازخورد انسانی می‌تواند میزان پیروی از دستورها و کیفیت تعامل را افزایش دهد.

سپس در ۳۰ نوامبر ۲۰۲۲ ChatGPT عرضه شد. از نظر علمی، بسیاری از اجزای اصلی آن از قبل وجود داشتند، اما از نظر اجتماعی ChatGPT یک نقطه عطف بود. رابط گفت‌وگویی پیچیدگی تعامل با مدل‌های بزرگ را از دید کاربر پنهان کرد. اکنون فردی که هیچ دانش برنامه‌نویسی نداشت می‌توانست با زبان روزمره از ماشین بخواهد مقاله‌ای را خلاصه کند، برنامه‌ای بنویسد، یک مفهوم علمی را توضیح دهد، متنی را ویرایش کند، ایده‌پردازی کند یا داده‌ای را تحلیل نماید. OpenAI در معرفی اولیه ChatGPT در نوامبر ۲۰۲۲ آن را مدلی گفت‌وگومحور مرتبط با InstructGPT معرفی کرد که برای پیروی از دستورهای کاربر طراحی شده بود.

از این نقطه، AI وارد مرحله‌ای شد که می‌توان آن را «دموکراتیزه شدن رابط هوش مصنوعی» نامید. تا پیش از آن، استفاده مؤثر از بسیاری از سیستم‌های AI نیازمند برنامه‌نویسی یا مهارت تخصصی بود. مدل‌های مکالمه‌ای زبان طبیعی را به رابطی عمومی تبدیل کردند. این امر اهمیت تاریخی دارد، زیرا زبان طبیعی ابزار اصلی ارتباط اجتماعی انسان است. اگر رایانه‌های شخصی رابط گرافیکی را به رابط عمومی انسان و کامپیوتر تبدیل کردند، مدل‌های زبانی بزرگ امکان آن را ایجاد کرده‌اند که زبان روزمره خود انسان به یکی از رابط‌های اصلی کنترل سیستم‌های دیجیتال تبدیل شود.

اما تحول در مدل‌های زبان متوقف نماند. مدل‌ها به تدریج «چندوجهی» یا Multimodal شدند. انسان جهان را فقط از طریق متن تجربه نمی‌کند؛ تصویر، صدا، زبان، حرکت و محیط فیزیکی همگی بخشی از شناخت هستند. مدل‌های چندوجهی نیز می‌توانند اطلاعات را از چند قالب دریافت یا تولید کنند. یک سیستم می‌تواند تصویر را تحلیل کند و درباره آن صحبت کند، یک نمودار را توضیح دهد، صدای انسان را دریافت کند، با او مکالمه کند، تصاویر تولید کند یا اطلاعات متنی و تصویری را با یکدیگر ترکیب نماید. به این ترتیب مرز سنتی میان «سیستم تشخیص تصویر»، «مدل زبان»، «سیستم صوتی» و «نرم‌افزار» در حال کم‌رنگ شدن است.

تحول مهم بعدی، برجسته شدن مدل‌های استدلالی بود. مدل‌های زبانی اولیه در تولید متن روان بسیار قوی بودند، اما در مسائل پیچیده چندمرحله‌ای، ریاضیات یا برنامه‌نویسی اغلب دچار خطا می‌شدند. نسل‌های جدیدتر AI بیشتر بر فرایند حل مسئله متمرکز شده‌اند: تجزیه یک هدف پیچیده به مراحل مختلف، بررسی راه‌حل‌های متفاوت، استفاده از ابزارهای بیرونی، اجرای محاسبات و ارزیابی نتیجه. گزارش AI Index 2026 دانشگاه Stanford نشان می‌دهد که پیشرفت در برخی آزمون‌های دشوار علمی، ریاضی، استدلال چندوجهی و برنامه‌نویسی در فاصله کوتاهی بسیار سریع بوده است. طبق این گزارش، برخی مدل‌های مرزی در شماری از آزمون‌ها به معیارهای انسانی رسیده یا از آنها عبور کرده‌اند و عملکرد روی بعضی بنچمارک‌های برنامه‌نویسی نیز در یک سال افزایش بسیار بزرگی داشته است.

در کنار مدل‌های استدلالی، مفهوم Agentic AI یا «هوش مصنوعی عامل‌محور» به یکی از مهم‌ترین جهت‌های توسعه تبدیل شده است. تفاوت میان chatbot و agent اساسی است. یک chatbot معمولاً سؤال دریافت می‌کند و پاسخ می‌دهد. اما یک عامل هوشمند می‌تواند هدفی دریافت کند، آن را به مجموعه‌ای از مراحل تبدیل کند، درباره ترتیب کار تصمیم بگیرد، از موتور جست‌وجو، پایگاه داده، نرم‌افزار، فایل یا کد استفاده کند، نتیجه یک مرحله را بررسی کند و بر اساس آن مرحله بعدی را انتخاب نماید. در چنین مدلی AI از یک «ماشین پاسخ‌دهنده» به نوعی «سامانه انجام وظیفه» تبدیل می‌شود. پژوهشگران Dartmouth نیز در سال ۲۰۲۶، ضمن بررسی Agentic AI، بر این نکته تأکید کرده‌اند که سیستم‌های عامل‌محور مرحله مهمی در توسعه AI هستند، هرچند همچنان با محدودیت‌هایی در قابلیت اعتماد و خودمختاری روبه‌رو هستند.

این تحول می‌تواند از نظر اقتصادی حتی عمیق‌تر از ظهور chatbotها باشد. هنگامی که AI فقط یک متن تولید می‌کند، ابزار افزایش بهره‌وری است؛ اما هنگامی که می‌تواند بخشی از یک فرایند تولید، تحقیق، برنامه‌نویسی، مدیریت، پشتیبانی یا تحلیل را از ابتدا تا انتها پیش ببرد، به یک جزء فعال در سازماندهی کار تبدیل می‌شود. بنابراین مسئله آینده AI تنها این نیست که ماشین «چه می‌داند»، بلکه این است که چه میزان از چرخه تبدیل هدف به تصمیم و تصمیم به عمل می‌تواند به سیستم‌های هوشمند واگذار شود.

تا سال ۲۰۲۶، هوش مصنوعی به مرحله‌ای رسیده است که توصیف آن صرفاً به عنوان یک رشته علوم کامپیوتر دیگر کافی نیست. این فناوری در حال تبدیل شدن به زیرساخت عمومی فعالیت‌های شناختی است. مدل‌های پیشرفته می‌توانند متن، تصویر، صدا و کد را پردازش کنند، در تحقیقات علمی کمک کنند، داده‌های بزرگ را تحلیل کنند، در تولید نرم‌افزار مشارکت کنند، مسائل ریاضی حل کنند، با ابزارهای خارجی ارتباط برقرار نمایند و در قالب agent وظایف چندمرحله‌ای را دنبال کنند. Stanford AI Index 2026 همچنین گزارش می‌کند که بیش از ۹۰ درصد مدل‌های برجسته مرزی سال ۲۰۲۵ توسط صنعت تولید شده‌اند. این واقعیت نشان می‌دهد مرکز ثقل پژوهش AI که زمانی عمدتاً در دانشگاه‌ها و آزمایشگاه‌های عمومی قرار داشت، در حوزه مدل‌های مرزی به شرکت‌هایی منتقل شده که توانایی تأمین هزینه عظیم تراشه، مراکز داده، برق، شبکه و نیروی متخصص را دارند.

در اینجا تاریخ تکنیکی هوش مصنوعی به تاریخ اقتصادی آن پیوند می‌خورد. AI جدید صرفاً یک الگوریتم نیست؛ حاصل ترکیب الگوریتم، داده، تراشه، انرژی، مراکز داده، شبکه ارتباطی و سرمایه عظیم است. به همین دلیل بحث درباره هوش مصنوعی امروز ناگزیر بحث درباره مالکیت زیرساخت‌های آن نیز هست. یک مدل پیشرفته ممکن است بر حجم عظیمی از دانش انباشته انسانی ــ کتاب‌ها، مقالات علمی، نرم‌افزارها، تصاویر، زبان و دیگر آثار فرهنگی ــ آموزش دیده باشد، اما توانایی ساخت و کنترل مدل‌های مرزی در اختیار شمار نسبتاً محدودی از نهادهای بزرگ اقتصادی قرار دارد. در این نقطه یک تضاد قابل توجه تاریخی شکل می‌گیرد: مواد اولیه معرفتی AI عمیقاً اجتماعی هستند، در حالی که ابزار پردازش و تبدیل آنها به قدرت اقتصادی می‌تواند بسیار متمرکز باشد.

همین مسئله به یکی از بزرگ‌ترین پرسش‌های اجتماعی عصر هوش مصنوعی منتهی می‌شود. پرسش آینده فقط این نیست که «ماشین چه کارهایی می‌تواند انجام دهد؟» بلکه باید پرسید چه کسی مالک ماشین، داده، مدل، تراشه و زیرساخت است؛ چه کسی اهداف سیستم را تعیین می‌کند؛ چه کسی حق دسترسی دارد؛ چه کسانی از افزایش بهره‌وری آن بهره‌مند می‌شوند؛ و هزینه‌های اجتماعی آن بر دوش چه کسانی قرار می‌گیرد.

تاریخ هوش مصنوعی از این زاویه بخشی از تاریخ نیروهای مولد نیز هست. ماشین بخار قدرت عضلانی انسان را چند برابر کرد. موتور الکتریکی و خط تولید امکان سازماندهی تولید انبوه را فراهم کردند. رایانه محاسبه و پردازش اطلاعات را خودکار کرد. اینترنت امکان انتقال و اشتراک تقریباً آنی اطلاعات را در مقیاس جهانی ایجاد نمود. هوش مصنوعی اکنون وارد قلمروی دیگری شده است: بخشی از کار شناختی.

کار شناختی شامل نوشتن، ترجمه، تحلیل، طراحی، برنامه‌نویسی، تشخیص الگو، برنامه‌ریزی، تحقیق و بخشی از تصمیم‌سازی است؛ فعالیت‌هایی که تا همین اواخر تصور می‌شدند به‌طور انحصاری یا تقریباً انحصاری به انسان تعلق دارند. هوش مصنوعی الزاماً تمام این فعالیت‌ها را جایگزین نمی‌کند، اما رابطه میان انسان و ابزار تولید را در آنها تغییر می‌دهد. همان‌گونه که ورود ماشین صنعتی الزاماً کار فیزیکی انسان را فوراً حذف نکرد اما ماهیت و سازمان آن را دگرگون کرد، AI نیز در حال تغییر ساختار کار فکری است.

با این حال باید از یک اشتباه مهم پرهیز کرد: هوش مصنوعی و آگاهی یک مفهوم نیستند. پیشرفت حیرت‌آور مدل‌های امروز نشان می‌دهد ماشین می‌تواند در برخی وظایف عملکردی را تولید کند که پیش‌تر نشانه هوش انسانی دانسته می‌شد، اما این امر به خودی خود نشان نمی‌دهد که ماشین دارای consciousness یا تجربه ذهنی است. یک مدل می‌تواند هزاران صفحه درباره درد، عشق، ترس یا مرگ تولید کند بدون آنکه شواهدی داشته باشیم که خود آن تجربیات را احساس می‌کند. می‌تواند درباره یک مفهوم استدلال کند بدون آنکه بتوانیم از این عملکرد نتیجه بگیریم دارای تجربه خودآگاهانه‌ای مشابه انسان است. بنابراین باید میان پردازش اطلاعات، یادگیری، هوش، شناخت و آگاهی تمایز قائل شد.

این تمایز از نظر آینده اجتماعی AI اهمیت زیادی دارد. ممکن است جامعه‌ای دارای بالاترین سطح هوش محاسباتی باشد، اما الزاماً جامعه‌ای آگاه‌تر نباشد. یک سیستم فوق‌العاده قدرتمند می‌تواند برای کشف دارو، آموزش، برنامه‌ریزی اقتصادی، کاهش اتلاف منابع یا گسترش دسترسی به دانش مورد استفاده قرار گیرد؛ همان سیستم یا فناوری مشابه می‌تواند در نظارت گسترده، دستکاری اطلاعات، جنگ، انحصار اقتصادی یا کنترل اجتماعی نیز به کار رود. فناوری به خودی خود جهت اخلاقی یا اجتماعی نهایی خود را تعیین نمی‌کند. این انسان‌ها، نهادها و ساختارهای اجتماعی‌اند که تعیین می‌کنند ظرفیت فنی چگونه مورد استفاده قرار گیرد.

از همین رو، تاریخ هوش مصنوعی را می‌توان در دو سطح خواند. در سطح نخست، یک تاریخ تکنیکی مشاهده می‌کنیم: منطق ریاضی به محاسبات ماشینی، محاسبات به هوش نمادین، هوش نمادین به یادگیری ماشین، یادگیری ماشین به شبکه‌های عمیق، شبکه‌های عمیق به Transformer، Transformer به مدل‌های بنیادین، مدل‌های بنیادین به هوش مولد و چندوجهی، و آنها به مدل‌های استدلالی و عامل‌های هوشمند منتهی شده‌اند.

اما در سطح عمیق‌تر، تحول دیگری رخ داده است: ابتدا انسان قواعد را به ماشین می‌داد؛ سپس داده در اختیار آن گذاشت تا الگو را یاد بگیرد؛ سپس ماشین توانست بازنمایی‌های پیچیده‌ای از اطلاعات ایجاد کند؛ مدل‌های بزرگ بخشی از دانش اجتماعی انباشته‌شده را در ساختارهای آماری خود جذب کردند؛ و اکنون سیستم‌های عامل‌محور می‌کوشند این دانش را به اقدام تبدیل کنند.

به بیان فشرده، مسیر تحول را می‌توان چنین دید:

قاعده → داده → اطلاعات → الگو → مدل → دانش عملیاتی → استدلال → اقدام

این زنجیره شاید از هر جدول زمانی دیگری بهتر توضیح دهد که چرا هوش مصنوعی امروز با فناوری‌های پیشین تفاوت دارد. ابزار دیگر صرفاً چیزی نیست که انسان از آن برای اجرای یک دستور ثابت استفاده کند؛ ابزار به تدریج در تفسیر هدف، یافتن مسیر و انتخاب بعضی اقدامات نیز مشارکت می‌کند.

در نتیجه، پرسشی که آلن تورینگ در سال ۱۹۵۰ مطرح کرد ــ آیا ماشین‌ها می‌توانند فکر کنند؟ ــ هنوز از بین نرفته است، اما در سال ۲۰۲۶ دیگر تنها پرسش مهم نیست. ماشین‌ها اکنون قادرند رفتارهایی تولید کنند که در بسیاری حوزه‌ها عملاً هوشمندانه تلقی می‌شوند. مسئله اجتماعی بزرگ‌تر این است که انسان با این ظرفیت جدید چه خواهد کرد؟

آیا هوش مصنوعی عمدتاً ابزاری برای تمرکز بیشتر ثروت و قدرت خواهد شد، یا می‌تواند به گسترش دانش و مشارکت اجتماعی کمک کند؟ آیا بهره‌وری حاصل از آن به کاهش ساعات کار و ارتقای کیفیت زندگی منجر خواهد شد، یا به حذف فرصت‌های شغلی بدون توزیع منافع؟ آیا داده و دانش اجتماعی به منبع عمومی توانمندسازی انسان تبدیل خواهند شد یا ماده اولیه انحصارات جدید؟ آیا AI ابزار شفافیت قدرت خواهد بود یا ابزار نظارت بر شهروند؟ و مهم‌تر از همه، آیا افزایش «هوش» ماشینی به افزایش «آگاهی» انسانی و اجتماعی خواهد انجامید؟

شاید همین‌جا باشد که تاریخ هوش مصنوعی با بحث بزرگ‌تری درباره نظم اجتماعی نوین در عصر آگاهی پیوند می‌خورد. انقلاب صنعتی پرسش مالکیت کارخانه و ابزار تولید را به مرکز تحولات اجتماعی آورد. انقلاب دیجیتال مسئله مالکیت اطلاعات و شبکه‌ها را برجسته کرد. انقلاب هوش مصنوعی اکنون مسئله مالکیت و کنترل زیرساخت‌هایی را مطرح می‌کند که نه فقط کالاهای مادی، بلکه دانش، تصمیم و فعالیت شناختی را پردازش می‌کنند.

از این منظر، AI را نمی‌توان فقط به‌عنوان یک اختراع فنی بررسی کرد. این فناوری در حال تبدیل شدن به یکی از نیروهای مولد تعیین‌کننده قرن بیست‌ویکم است. همان‌گونه که فناوری‌های گذشته ساختار اقتصاد و جامعه را تغییر دادند، هوش مصنوعی نیز احتمالاً ساختار کار، آموزش، مالکیت، قدرت، سیاست و حتی تعریف ما از مهارت انسانی را دگرگون خواهد کرد.

بنابراین مهم‌ترین فصل تاریخ هوش مصنوعی شاید هنوز نوشته نشده باشد. نخستین فصل آن درباره این بود که آیا می‌توان ماشین را وادار به محاسبه کرد. فصل بعد پرسید آیا ماشین می‌تواند استدلال کند. سپس مسئله یادگیری، دیدن، شنیدن و فهم زبان مطرح شد. امروز سؤال به سمت توانایی ماشین برای حل مسئله و عمل نسبتاً مستقل حرکت کرده است. اما فصل آینده بیش از آنکه مسئله‌ای درباره ماشین باشد، مسئله‌ای درباره جامعه خواهد بود.

تا امروز پرسش اصلی دانشمندان این بوده است:

ماشین تا چه اندازه می‌تواند هوشمند شود؟

پرسش تعیین‌کننده عصر آینده احتمالاً این خواهد بود:

جامعه انسانی با این هوش چه خواهد کرد؟

پاسخ به پرسش نخست عمدتاً در آزمایشگاه‌ها، دانشگاه‌ها و مراکز داده نوشته شده است. پاسخ به پرسش دوم را دیگر مهندسان به تنهایی نمی‌توانند تعیین کنند. این پاسخ به اقتصاد، سیاست، اخلاق، حقوق، فرهنگ، مالکیت، مشارکت اجتماعی و در نهایت به سطح آگاهی جامعه وابسته خواهد بود.

در همین معنا، تاریخ هوش مصنوعی دیگر فقط تاریخ ماشین‌های هوشمند نیست؛ به تدریج به بخشی از تاریخ تحول خود جامعه انسانی تبدیل شده است.

منابع و مطالعات بیشتر

  1. Turing, A. M. (1950). Computing Machinery and Intelligence. Mind, 59(236), 433–460. مقاله بنیادی تورینگ که مسئله «آیا ماشین‌ها می‌توانند فکر کنند؟» و بازی تقلید را مطرح کرد.
  2. McCulloch, W. S., & Pitts, W. (1943). A Logical Calculus of the Ideas Immanent in Nervous Activity. Bulletin of Mathematical Biophysics, 5, 115–133. یکی از نخستین مدل‌های رسمی نورون و شبکه عصبی مصنوعی.
  3. McCarthy, J., Minsky, M., Rochester, N., & Shannon, C. (1955). A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence. پیشنهاد تاریخی پروژه‌ای که در تابستان ۱۹۵۶ به شکل‌گیری رسمی رشته هوش مصنوعی انجامید.
  4. Rosenblatt, F. (1958). The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain. Psychological Review, 65(6), 386–408. از متون بنیادی شبکه‌های عصبی یادگیرنده.
  5. Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323, 533–536. از مهم‌ترین مقالات در احیای شبکه‌های عصبی چندلایه و backpropagation.
  6. Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet Classification with Deep Convolutional Neural Networks. Advances in Neural Information Processing Systems. مقاله AlexNet و یکی از نقاط عطف انقلاب یادگیری عمیق.
  7. Silver, D., et al. (2017). Mastering the Game of Go without Human Knowledge. Nature. پژوهش AlphaGo Zero درباره یادگیری تقویتی و self-play.
  8. Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention Is All You Need. مقاله معرفی معماری Transformer که بنیان بخش بزرگی از مدل‌های زبانی و مولد امروزی شد.
  9. Brown, T. B., et al. (2020). Language Models are Few-Shot Learners. پژوهش GPT-3 درباره رابطه افزایش مقیاس مدل‌های زبانی با قابلیت‌های عمومی‌تر و few-shot learning.
  10. Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback. پژوهش InstructGPT و استفاده از بازخورد انسانی برای همسوسازی رفتار مدل‌های زبان.
  11. OpenAI (2022). Introducing ChatGPT. اعلام انتشار عمومی ChatGPT در ۳۰ نوامبر ۲۰۲۲.
  12. Stanford Institute for Human-Centered Artificial Intelligence (2026). AI Index Report 2026. گزارش جامع درباره وضعیت فنی، اقتصادی، علمی و اجتماعی هوش مصنوعی تا سال ۲۰۲۶.

مطالعه نسخهٔ انگلیسی ↓


This article traces the history of artificial intelligence from its theoretical foundations through developments in 2026.

A History of Artificial Intelligence: From the Dream of a Thinking Machine to Generative AI, Reasoning Models, and Intelligent Agents

Artificial intelligence has entered everyday life, the economy, education, scientific research, industry, media, and politics so rapidly that it can seem like a phenomenon of only the last few years. The release of ChatGPT at the end of 2022, the spread of generative image and video models, the development of multimodal systems, and then the arrival of systems able to reason, write programs, use tools, and independently carry out parts of a workflow have unquestionably accelerated this transformation. Yet what we now call artificial intelligence is the result of more than eight decades of scientific development, with theoretical roots that precede modern electronic computers. The history of AI is the history of the convergence of mathematical logic, neuroscience, computation theory, statistics, computer engineering, linguistics, cognitive psychology, and, more recently, data science and the digital economy.

At the center of this history lies an old question: does what humans call “thinking” possess a structure that can be converted into computable operations? If so, can a machine do more than calculate numbers—can it recognize patterns, understand language, learn from experience, make decisions, and display behavior we would call intelligent?

The answers offered over the past eighty years have changed. Early researchers thought intelligence could be expressed as logical rules supplied to a machine. It later became clear that many dimensions of human intelligence cannot easily be written as explicit rules. Attention therefore shifted from programming the rules of intelligence to learning from data. Deep neural networks then made it possible to extract highly complex patterns. The Transformer architecture of 2017 opened the way to foundation models and very large language models. From the early 2020s, AI moved from being primarily a specialist tool toward becoming a general technology through which people could communicate in natural language. The current transition is moving it from “a machine that answers” toward “a system that can act to achieve a goal.”

To understand this path, we must begin before the term artificial intelligence existed. In the 1930s, British mathematician Alan Turing posed one of computer science’s most fundamental questions: what, in principle, is computable? His theoretical model, later called the Turing machine, gave computation a precise mathematical definition. Its importance for AI was that it established a theoretical framework for operations performed by a machine: if a process can be expressed as a definite sequence of operations, a general computing machine can in principle execute it.

In 1943, Warren McCulloch, a neuroscientist and psychiatrist, and Walter Pitts, a young logician, published “A Logical Calculus of the Ideas Immanent in Nervous Activity.” They proposed an abstract model of a neuron in which simple computational units could switch on or off according to their inputs and connect into networks. Their model differed greatly from present-day neural networks, but it was historically foundational because it established a formal relationship between neural activity and computational logic. It introduced the possibility that some functions of the nervous system might be modeled by networks of computational elements.

In 1950, Turing moved the question forward in “Computing Machinery and Intelligence,” published in Mind. Instead of becoming trapped in philosophical definitions of “machine” and “thought,” he proposed evaluating behavior through the “imitation game,” later known as the Turing test. In its simplest interpretation, if a human conversing through text cannot reliably determine whether the other participant is a human or a machine, the machine has displayed behavior that may practically be called intelligent. Turing’s contribution was not merely a test. He examined many objections to machine intelligence that remain familiar today and established the relationship between computation and intelligence as a serious research problem.

The term “artificial intelligence,” however, had not yet been coined. The event commonly treated as the formal birth of AI as a scientific field was the Dartmouth Summer Research Project of 1956. Its 1955 proposal was prepared by John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon. McCarthy used “artificial intelligence” to name the new field. Their conjecture was ambitious: every aspect of learning or intelligence could, in principle, be described precisely enough for a machine to simulate it. The 1956 gathering brought together researchers who became major founders of the field, making 1956 its conventional date of birth.

The 1950s and 1960s were years of extraordinary optimism. Early successes encouraged the belief that genuinely intelligent machines might be close. Allen Newell, Herbert Simon, and Cliff Shaw developed the Logic Theorist, which proved some theorems in mathematical logic. The General Problem Solver followed, attempting to provide general methods for solving problems. These systems embodied what became known as symbolic AI: the world could be represented through symbolic objects, concepts, and relations, while reasoning could be implemented as logical rules. Given sufficient knowledge and rules of inference, a machine could draw conclusions.

Symbolic AI worked well on constrained problems but faced a fundamental limitation. People perform many everyday activities without possessing an explicit rulebook. Recognizing a friend in the street, understanding a joke, detecting a speaker’s tone, or distinguishing thousands of objects cannot easily be converted into endless “if–then” statements. Much of human intelligence arises through experience and pattern learning that people themselves cannot always express as rules.

Alongside symbolic AI, another path was developing. In the late 1950s, Frank Rosenblatt created the Perceptron, a neuron-inspired system able to adjust its connection weights from training examples. Instead of requiring a programmer to determine every decision rule in advance, it could learn relationships from sample data. This became the crucial distinction between traditional programming and machine learning: rather than telling the computer exactly what to do for every input, one supplies examples from which it discovers a decision pattern. Rosenblatt’s 1958 paper in Psychological Review became a direct ancestor of modern neural networks.

Hardware and theoretical limitations soon became obvious. Computers had little processing power or memory, and digital data was scarce. Researchers also discovered that tasks easy for people could be extremely difficult for machines. A young child can learn to distinguish cats from dogs from only a few examples, whereas machines historically needed vast datasets. Translation, speech understanding, object recognition, and common-sense reasoning about real situations proved far harder than proving narrowly defined logical theorems.

During the 1970s, the gap between early promises and practical achievements produced cuts in funding and support. The period became known as an “AI winter.” AI has experienced more than one such cycle of excitement and disappointment. These cycles demonstrate that its progress has been neither linear nor inevitable; each advance has required a particular combination of scientific, hardware, and economic developments.

AI gained attention again in the 1980s through expert systems. These programs stored the knowledge of specialists as large collections of rules. A medical system could assess possible diagnoses from symptoms and test results, while an industrial system could recommend equipment configurations. MYCIN in medicine and XCON in computing became well-known examples. Expert systems showed that AI could create economic value in restricted domains, but maintaining thousands of rules became costly. Knowledge had to be extracted from experts, made explicit, recorded, and continually updated. The real world was too complex and dynamic to remain confined within fixed rules.

The same decade brought an important revival of neural networks. In 1986, David Rumelhart, Geoffrey Hinton, and Ronald Williams published “Learning Representations by Back-Propagating Errors” in Nature. Backpropagation measures the difference between a network’s output and the desired output, then carries the error backward through the layers so connection weights can be adjusted to reduce it. This method later became a principal foundation for training deep neural networks.

Neural networks nevertheless remained limited until three elements matured together: data, computing power, and algorithms. From the 1990s, and especially during the 2000s, the internet, smartphones, e-commerce, social media, digital systems, and sensors generated enormous datasets. Processor performance multiplied, and graphics processing units—originally designed for computer graphics—proved highly effective for the parallel operations required by neural networks. Algorithms, network-training methods, large datasets, and software tools also improved.

The meaning of AI gradually changed. Rather than having researchers directly supply all knowledge, machines extracted statistical relationships from data. In a traditional spam filter, a programmer might write rules about particular words or senders. In machine learning, thousands or millions of emails labeled “spam” or “normal” are supplied, and the model learns the relevant combination of features. Machine knowledge thus shifted from something directly written by a programmer to something extracted from data.

In 1997, IBM’s Deep Blue defeated world chess champion Garry Kasparov in a formal match. Chess had long symbolized strategic thought and human intelligence, so the victory had enormous public significance. Yet Deep Blue was fundamentally unlike today’s general models. It was designed specifically for chess and relied on extensive search, position evaluation, and specialized knowledge. It showed that a machine could surpass humans in one bounded domain, not that it possessed flexible general intelligence.

The next great turning point arrived in 2012. Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton trained a deep neural network for the ImageNet image-recognition competition. AlexNet, trained on more than a million images, dramatically reduced error compared with earlier methods. It demonstrated that deep neural networks, given enough data and computation, could learn complex visual features for themselves. This marked the practical beginning of the deep-learning explosion of the 2010s, after which deep networks rapidly became dominant in vision, speech, translation, language processing, medicine, and many other fields.

Deep learning uses neural networks with many layers, each producing a more complex representation of data. In image recognition, early layers may identify edges; intermediate layers learn shapes or textures; deeper layers recognize structures such as eyes, faces, animals, or vehicles. The engineer no longer needs to define all relevant features manually—the network extracts them from data.

Four years after AlexNet, another symbolic event changed public perceptions. In March 2016, DeepMind’s AlphaGo defeated Lee Sedol, one of the world’s leading Go players, four games to one. Because Go has an astronomical number of possible states, it had long been expected to resist machines much longer than chess. AlphaGo combined neural networks, reinforcement learning, and tree search, showing that learning and computation could reach solutions beyond brute-force enumeration.

AlphaGo Zero and AlphaZero went further. Through reinforcement learning and self-play, systems learned powerful strategies without relying on vast archives of human games. Starting with only the rules, AlphaGo Zero improved its network by playing against itself and effectively became its own teacher. Machine learning had entered a new phase: the machine could learn not only from human-generated data but also from experience it generated itself.

The development that led most directly to today’s large language models came in 2017, when Google researchers published “Attention Is All You Need” and introduced the Transformer. Earlier language systems often relied on recurrent architectures that struggled with long sequences. The Transformer used attention, allowing a model to assess how every part of a text relates to other parts while enabling far greater parallelism during training. Initially demonstrated in machine translation, the architecture later became the foundation of most large language models, multimodal models, and contemporary generative systems.

After the Transformer, the idea of a foundation model grew in importance. Instead of building separate small models for translation, summarization, and question answering, one very large model could be trained on massive bodies of information and then used for many tasks. GPT-2 in 2019 showed that a large language model trained on broad text could produce coherent passages and display abilities in answering questions, summarizing, and translating without separate task-specific training.

GPT-3, introduced in 2020 with 175 billion parameters, created another leap. “Language Models Are Few-Shot Learners” showed that greater scale could expand a model’s ability to perform many tasks from a few examples or natural-language instructions. This changed the human–machine interface. In conventional software, changing behavior requires changing code; with large language models, users could increasingly guide behavior through ordinary language. The prompt became a new interface between human intention and computation.

Another technical and social development concerned training models to respond helpfully to people. Predicting the next token alone does not guarantee that a system will understand a request or provide a useful response. Reinforcement learning from human feedback (RLHF) allowed models to learn from human preferences about answer quality. InstructGPT became an important example; research published in 2022 showed that human-feedback tuning could improve instruction-following and interaction.

ChatGPT was released on November 30, 2022. Most of its scientific components already existed, but socially it was a turning point. Its conversational interface concealed the complexity of large models. A person with no programming knowledge could use everyday language to request a summary, a program, an explanation, an edit, ideas, or data analysis. AI entered what might be called the democratization of the AI interface. Just as graphical interfaces made personal computing broadly accessible, conversational models made ordinary language a principal way to direct digital systems.

Models then became increasingly multimodal. Humans do not experience the world through text alone: image, sound, language, movement, and physical surroundings all participate in cognition. Multimodal systems can receive or produce information in several forms. They can analyze an image and discuss it, explain a chart, hear and answer speech, produce images, or combine textual and visual information. The boundaries among image recognition, language models, audio systems, and software have consequently begun to blur.

The next major direction was the rise of reasoning models. Early language models produced fluent text but often failed on complex, multistep problems in mathematics or programming. Newer AI systems place greater emphasis on problem solving: decomposing a goal into steps, considering alternatives, using external tools, executing calculations, and evaluating results. Stanford’s 2026 AI Index reports rapid gains in difficult scientific, mathematical, multimodal-reasoning, and coding evaluations. Some frontier models have reached or surpassed human benchmarks on selected tests, while performance on certain software-engineering benchmarks has risen sharply within a short period.

Alongside reasoning models, agentic AI has become a major development path. The difference between a chatbot and an agent is fundamental. A chatbot normally receives a question and returns an answer. An intelligent agent can receive a goal, break it into stages, decide on an order of work, use search engines, databases, applications, files, or code, inspect the result of one stage, and choose the next. AI is changing from an answering machine into a task-performing system. Yet agentic systems still face important limits in reliability, oversight, and autonomy.

This transition may prove even more economically consequential than chatbots. When AI only produces text, it is a productivity tool. When it can conduct part of a production, research, programming, administrative, support, or analytical process from beginning to end, it becomes an active component in organizing work. The future question is therefore not only what a machine knows, but how much of the cycle from goal to decision and from decision to action may be entrusted to intelligent systems.

By 2026, describing AI as merely another branch of computer science is no longer adequate. It is becoming a general infrastructure for cognitive activity. Advanced models can process text, images, audio, and code; assist scientific research; analyze large datasets; participate in software development; solve mathematical problems; connect to external tools; and pursue multistep tasks as agents. Stanford’s 2026 AI Index also reports that industry produced more than 90 percent of notable frontier models in 2025. The center of gravity in frontier AI has shifted from universities and public laboratories toward companies able to finance enormous requirements for chips, data centers, electricity, networks, and specialized labor.

Here the technical history of AI joins its economic history. Contemporary AI is not simply an algorithm; it is the combination of algorithms, data, chips, energy, data centers, communications networks, and immense capital. Discussion of AI therefore necessarily includes ownership of its infrastructure. Advanced models may be trained on humanity’s accumulated knowledge—books, research, software, images, languages, and cultural works—while the capacity to build and control frontier systems remains concentrated in relatively few large institutions. A historical contradiction emerges: AI’s intellectual raw material is profoundly social, while the means of processing it into economic power may be highly concentrated.

The question of the future is not only what machines can do. We must ask who owns the machines, data, models, chips, and infrastructure; who determines system objectives; who has access; who receives the benefits of increased productivity; and who bears the social costs.

From this perspective, AI history is also part of the history of productive forces. The steam engine multiplied human muscular power. Electric motors and assembly lines enabled mass production. Computers automated calculation and information processing. The internet enabled near-instant transmission and sharing of information at global scale. Artificial intelligence is now entering another realm: part of cognitive labor.

Cognitive labor includes writing, translation, analysis, design, programming, pattern recognition, planning, research, and some decision support—activities until recently regarded as exclusively or almost exclusively human. AI will not necessarily replace all of them, but it changes the relationship between people and their tools of production. Industrial machines did not instantly eliminate physical labor, but transformed its character and organization; AI is likewise changing the structure of intellectual work.

One important mistake must be avoided: intelligence and consciousness are not the same. The remarkable performance of contemporary models shows that machines can produce outputs once treated as signs of human intelligence, but it does not by itself show that machines possess consciousness or subjective experience. A model may write thousands of pages about pain, love, fear, or death without evidence that it feels any of them. It may reason about a concept without possessing self-aware experience comparable to that of a human being. We must distinguish information processing, learning, intelligence, cognition, and consciousness.

This distinction matters socially. A society may possess an unprecedented level of computational intelligence without becoming more conscious. A powerful system can support drug discovery, education, economic planning, the reduction of waste, and wider access to knowledge. The same or similar technology can be used for mass surveillance, information manipulation, warfare, monopoly, or social control. Technology does not determine its own ethical and social direction. Human beings, institutions, and social structures determine how technical capacity is used.

AI history can therefore be read at two levels. At the technical level, mathematical logic led to machine computation; computation to symbolic AI; symbolic AI to machine learning; machine learning to deep networks; deep networks to the Transformer; the Transformer to foundation models; and foundation models to generative and multimodal AI, reasoning systems, and intelligent agents.

At a deeper level, humans first gave machines rules; then supplied data from which patterns could be learned; machines created increasingly complex representations of information; large models absorbed part of society’s accumulated knowledge into statistical structures; and agentic systems now attempt to turn that knowledge into action.

Rule → Data → Information → Pattern → Model → Operational Knowledge → Reasoning → Action

This chain explains why contemporary AI differs from earlier technology. A tool is no longer used only to execute a fixed instruction. It increasingly participates in interpreting the objective, finding a path, and selecting some actions.

The question Turing asked in 1950—can machines think?—has not disappeared, but by 2026 it is no longer the only important question. Machines can already produce behavior treated as intelligent across many domains. The larger social question is what humanity will do with this new capacity.

Will AI become primarily an instrument for concentrating wealth and power, or can it expand knowledge and social participation? Will its productivity gains reduce working time and improve life, or eliminate jobs without distributing the benefits? Will social knowledge and data become public resources for human empowerment or raw material for new monopolies? Will AI make power more transparent, or citizens more observable? Most importantly, will greater machine “intelligence” contribute to greater human and social “consciousness”?

This is where the history of AI connects to the broader discussion of A New Social Order in the Age of Consciousness. The Industrial Revolution placed ownership of factories and productive tools at the center of social conflict. The digital revolution foregrounded ownership of information and networks. The AI revolution now raises the issue of owning and controlling infrastructures that process not only material goods, but knowledge, decisions, and cognitive activity.

AI cannot be understood solely as a technical invention. It is becoming one of the decisive productive forces of the twenty-first century. As earlier technologies transformed economy and society, AI will likely reshape work, education, ownership, power, politics, and even our definition of human skill.

The most important chapter in AI history may therefore remain unwritten. Its first chapter asked whether machines could calculate. The next asked whether they could reason. Later came learning, vision, hearing, and language. Today the question is moving toward problem solving and relatively autonomous action. Yet the next chapter will be less a question about machines than a question about society.

Until now, scientists have principally asked: How intelligent can a machine become?

The decisive question of the coming age may be: What will human society do with this intelligence?

The answer to the first question has largely been written in laboratories, universities, and data centers. Engineers alone cannot determine the answer to the second. It will depend on economics, politics, ethics, law, culture, ownership, social participation, and ultimately the level of society’s consciousness.

In this sense, the history of artificial intelligence is no longer only the history of intelligent machines. It is becoming part of the history of the transformation of human society itself.

References and Further Reading

  1. Turing, A. M. (1950). Computing Machinery and Intelligence. Mind, 59(236), 433–460.
  2. McCulloch, W. S., & Pitts, W. (1943). A Logical Calculus of the Ideas Immanent in Nervous Activity. Bulletin of Mathematical Biophysics, 5, 115–133.
  3. McCarthy, J., Minsky, M., Rochester, N., & Shannon, C. (1955). A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence.
  4. Rosenblatt, F. (1958). The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain. Psychological Review, 65(6), 386–408.
  5. Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning Representations by Back-Propagating Errors. Nature, 323, 533–536.
  6. Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet Classification with Deep Convolutional Neural Networks. Advances in Neural Information Processing Systems.
  7. Silver, D., et al. (2017). Mastering the Game of Go without Human Knowledge. Nature.
  8. Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention Is All You Need.
  9. Brown, T. B., et al. (2020). Language Models Are Few-Shot Learners.
  10. Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback.
  11. OpenAI (2022). Introducing ChatGPT.
  12. Stanford Institute for Human-Centered Artificial Intelligence (2026). AI Index Report 2026.

Read the Persian version ↑