creator cover Artem X
Artem X

Artem X 

ai-researcher

1subscriber

8posts

goals3
$0 of $275 raised
subscription to claude max 20x
$0 of $125 raised
Vast.ai GPU rentals
$0 of $872 raised
ML server: - 2× NVIDIA Tesla P100 16 GB PCIe GPUs - Dual X99 / C612 with 2× PCIe x16 - 2× Intel Xeon E5 v3/v4 CPUs - 32 GB DDR4 ECC - 1000 W supply

About

AI researcher, mostly thinking about transformer-based language models — how they work and how to improve them. But my main job is working with microcontrollers.
Habr https://habr.com/ru/users/Imperius14/
Twitter (X): https://x.com/vla3728419
Codeberg: https://codeberg.org/imperius
Hugging-Face: https://huggingface.co/Imperius
Medium: https://medium.com/@artem-x
Github: https://github.com/artem-x-meta

Building the Meta-Spider framework on top of meta-attention

This is a direct sequel to the article “meta-attention is all you need”. Reading it first is recommended but not required — a tour of the architecture is included below.
The first article described the idea in detail and published the experiment sources, but “sources” there meant simply all the code that had been written, with no system to it — including several different implementations of the same components.
So here is a framework with a ready-made toolkit you can try out on LLMs, including in agentic scenarios.
Ready-to-use, lightweight trained wrappers are also provided — one tiny model (Qwen-3.5–4B) and one mid-size one (Granite 3.3 8B). All of them can be run through llama.cpp.
https://medium.com/@artem-x/building-the-meta-spider-framework-on-top-of-meta-attention-476eb95bb85b?postPublishedType=initial

How I Grew a Digital Homunculus and Became a Neuro-Punk

Why? To create Skynet, of course.
Well, also because I wanted to understand, in detail, what this field that fascinates me so much is breathing with right now. And the best way to understand something is to try to explain it to someone else.
Besides that, I want to move into deep learning professionally, and publishing my interesting projects on the internet seems like the fastest way to get noticed.
Personally, I enjoyed the process a lot, and I invite Habr readers to dive into this small journey with me.
Links to the dataset, weights, and code are attached at the end of the article. The dataset and weights are on Hugging Face; the codebase is on Codeberg, a GitHub-like platform with a similar workflow.
Let's go.
https://dev.to/imperius_903049e65aa91ec5/how-i-grew-a-digital-homunculus-and-became-a-neuro-punk-19de

How I Loaded a Compact Open LLM Into a Robot and Told It to Walk (and Grab Things)

I will describe how I trained Google's 270M-parameter Gemma-3 language model to control a tracked robot with a robot arm in the MuJoCo environment, using natural-language commands from a human.
It can move freely around the map, go forward and backward, turn left and right, grab objects, and put them down.
https://dev.to/imperius_903049e65aa91ec5/how-i-loaded-a-compact-open-llm-into-a-robot-and-told-it-to-walk-and-grab-things-4h40

Terminator Is Still the Most Technically Accurate Depiction of AI, While Detroit: Become Human Is Science Fantasy

In this short essay, I want to reflect on how AI is depicted in fiction, or more specifically, on intelligent machines capable of solving all intellectual tasks at a human level or better.
Why are the first two Terminator films still the most realistic depiction of AI in fiction? What does James Cameron's technical background have to do with it? And why are intelligent computers almost always portrayed as "silicon humans"?
https://dev.to/imperius_903049e65aa91ec5/terminator-is-still-the-most-technically-accurate-depiction-of-ai-while-detroit-become-human-is-3jg7
Subscription levels0
No subscription levels
Go up