AI & ML interests

None defined yet.

Recent Activity

Hoglet-33Β 
posted an update 4 days ago
view post
Post
6174
Hey everyone! I got sidetracked from my main projects and decided to test out the BananaAll app and see if I could make a small model not regress too much during SFT. Here is what happened:

The base model I chose was BananaMind/BananaMind-2.1-Pico-Preview, and the dataset I used was SupraLabs/SupraThink-Dataset-500x

I trained for 5 whole steps using a LoRA adapter.

Results:
A model that scores better on some benchmarks and worse on others, and still lacks most general capabilities.

You can find the model here: Hoglet-33/Hogleto

Credits:

- Thank you to @Banaxi-Tech for the BananaAll app (works perfectly on Windows and CPU)
- GPT-6 Sol for knowing how to merge some confusing files created by the app
- Myself for the idea
- Someone else somewhere who might have contributed to some of my ideas and might in the future
- And readers like you!
  • 6 replies
Β·
Hoglet-33Β 
posted an update 7 days ago
view post
Post
2367
Everything going on here at basically AI:

1. Pebble 1.5

We're working on Pebble 1.5. Here's what we know so far:

- They will be better than the last generation. 99.99% certain.
- Expanded context lengths of at least 16,384 tokens, with the flagship potentially reaching 32,768.
- A Mamba3-based architecture with some other new architectural designs we're experimenting with.
- Native CPU compatibility β€” something we failed at with the last generation.
- Natively multilingual and multimodal???

2. SmolCodeBench

A code benchmark designed specifically for small models, because there really isn't a good one right now.

3. SENTRY

VOID is working on something called SENTRY β€” System for Evaluating Neural Threats, Responses, and Yields.

More on that soon.

4. basically OS

It's an operating system/app/harness. We're still deciding.

5. Finances

Trying to balance the finances after purchasing a Hugging Face Pro subscription.

Follow us for updates:
@Hoglet-33
basically-ai

basically-experimental

void-research
  • 2 replies
Β·
FlameF0XΒ 
posted an update 10 days ago
view post
Post
4578
Personally, I don't think bot accounts on Hugging Face are a good thing, as we don't know how many accounts are run by automated systems versus how many actual users there are. Dead Internet theory is already a thing.

To clarify I am taking about LLM powered bot accounts and NOT rule base once like @parquet-converter or others.

also I'd like to talk with HUMANS not a machine so I'm going to hide messages from bots.
  • 24 replies
Β·
Hoglet-33Β 
updated a Space 14 days ago
Hoglet-33Β 
posted an update 14 days ago
view post
Post
6317
Introducing VOID. A new research branch of basically AI.

VOID β€” Verification of Objectives, Intentions, and Deception.

We study what lies beneath the surface: objectives, intentions, and the possibility of deception in AI systems.

There isn't much to see yet.

That will change.

Follow us for updates:

@Hoglet-33
void-research

basically-ai
  • 7 replies
Β·
i64systemsΒ 
posted an update 24 days ago
view post
Post
89
we taught a small model to predict how simulated fluid would move eight steps into the future. it learned from examples, then succeeded on examples withheld from training.
two things mattered:
- its memory helped. the version with an internal memory produced about 41–45% less prediction error than the comparison model using recent observations.
- the learning was repeatable. restarting training reproduced every recorded update and saved checkpoint exactly. resuming halfway through also produced the same ending.
that gives us working evidence that bf16 gpu training can learn and remain exactly repeatable on this 3090 ti and software setup.

thank you for your time.
i64systemsΒ 
posted an update 25 days ago
view post
Post
101
little bob’s updated weights now change his actions in doom. all three fresh replays matched exactly

yes i am training my bf16/fp32 non-language models weights with replayable inference *in doom*

Hoglet-33Β 
posted an update 25 days ago
view post
Post
3234
Today, we planned to release Pebble-50M and Pebble-50M-Chat to the world. Unfortunately, due to a few issues, that didn't go quite as planned.

What happened:
- Some data and benchmark results were lost or corrupted
- The models performed worse on benchmarks than our other Pebble models

Despite that, you can still find both models here:
Pebble-50M-beta: basically-experimental/Pebble-50M-beta
Pebble-50M-Chat-beta: basically-experimental/Pebble-50M-Chat-beta

There are still some interesting improvements in these models:
- Compatible with non-CUDA devices
- Vocabulary increased to 16K tokens
- Context length increased to 16K tokens

For now, there won't be any more Pebble releases for a while. We're going to take some time to experiment with other approaches and hopefully make the next generation a monumental leap over this one.

Follow for updates:
@Hoglet-33
basically-ai

basically-experimental
  • 4 replies
Β·
FlameF0XΒ 
posted an update 26 days ago
view post
Post
3012
Hello HuggingFace! (UPDATE)

I tested the current architect of FWKV/Myosotis-1-base (that beinng FWKV) @ different sized and sequence lengths among RWKV and Transformer architecture. I did not include Mamba since that would require a costume kernel.

Note: the evaluation might not be accurate.
CSV avalible @ FlameF0X/evals

FlameF0X/cpu-lm-benchmark
  • 11 replies
Β·
i64systemsΒ 
posted an update 27 days ago
view post
Post
77
NOW UPDATED ❗❗

openbob 5.0.1πŸ­πŸ”Ž

An integer execution method for reproducible inference from publicly available model weights, demonstrated on Qwen3-4B. Journaled bytes and all.

Keep an eye out for the gpt-oss-120B on the 24gb GPU- deterministically.
We make AI models do the same things every time!β›“οΈπŸ˜ˆ
i64systems/Qwen3-4B-openbob-i8
FlameF0XΒ 
posted an update 28 days ago
Hoglet-33Β 
posted an update 29 days ago
view post
Post
3458
Pebble-25M and Pebble-25M-Chat are out now!

We’re excited to release Pebble-25M and Pebble-25M-Chat!

Both models use our 3:1 Mamba2/Transformer hybrid architecture and were pretrained on 25B tokens. Pebble-25M-Chat was then further fine-tuned on an additional 250M tokens from smol-smoltalk, following the same approach used for the Pebble-10M models.

We hope you enjoy experimenting with them!

Pebble-50M is coming in a few days.

Models
Pebble-25M: basically-ai/Pebble-25M
Pebble-25M-Chat: basically-ai/Pebble-25M-Chat
Pebble-10M GGUFs

In case you missed it, our friend @ContextReq made GGUF versions of the Pebble-10M models:

https://huggingface.co/ContextReq/Pebble-10M-GGUF
https://huggingface.co/ContextReq/Pebble-10M-Chat-GGUF

Follow us if you don’t want to miss future releases and updates!

@Hoglet-33
basically-ai

basically-experimental
  • 7 replies
Β·
i64systemsΒ 
posted an update 30 days ago
view post
Post
99
remat is no longer a one-model claim!!proved it on Qwen3-30B-A3B, K=32 of 128 experts resident, output task byte-identical to the full reference, zero bytes different *in bf16*πŸ₯°πŸ₯° GPU comes next😈

  • 6 replies
Β·
Hoglet-33Β 
published a Space about 1 month ago
FlameF0XΒ 
posted an update about 1 month ago
view post
Post
2634
Hello HuggingFace!

I would like to share a small preview of a language model that I have been experiencing for a while. In the first attached image there is a small sample of the Myosotis-1, an attempt to make a small 100m parameter flagship model that is built on my bizarre architecture that is somewhat similar to an S4/S5 model with WKV added. I call it FWKV (Feed-Forward WKV). The model is currently still in training because of the nature of RNN-like models. It can also be seen that the model has insane prompt processing and token generation speed (evaluation done on a 2x Titan XP); even for its small size, some similar Transformer models do struggle to get the same results without custom kernels (some Transformer models can achieve this level of throughput on cheap hardware).

In the second image, you can see checkpoint 20k of the model in its next token prediction state (this means it can't chat), ranking in the top 100 on AxiomicLabs/Open_SLM_Leaderboard (the results have not been submitted since the model is not done training).

- Why not just use Transformers?
Have you seen any pure non-Transformers SLMs besides RWKV and Mamba?

- Should you expect this project to become the next LFM or another very fast language model thing on some Raspberry Pi?
No, the model is still an experiment; it's very sensible and prone to collapse (by the time of this post, it can be seen in the 1st image).

- Should you use it?
Maybe not yet; the architecture itself is still very "naive"β€”that's how I could call it at its current level. If you just want to play with it and see what you could do or how fast the model is on your hardware, then you can do it.

Once the training is finished and I feel satisfied with the model next token prediction (the base model) and "assistants" (the instruction-tuned model) capabilities, I will make open weights at
FWKV
with full support of the πŸ€— Transformers.
  • 10 replies
Β·
Hoglet-33Β 
posted an update about 1 month ago
view post
Post
3323
Pebble 10M and Pebble 10M Chat are now released!

Both models use our Mamba/Transformer 3:1 hybrid architecture and were pretrained on 25 billion tokens.

Pebble 10M Chat was additionally fine-tuned on 250 million tokens of Smol-SmolTalk to improve its conversational capabilities.

You can find them here:

- basically-ai/Pebble-10M
- basically-ai/Pebble-10M-Chat

We hope you enjoy using them. The rest of the Pebble family will be released soon.

Follow for more:
@Hoglet-33
basically-ai
  • 32 replies
Β·
Hoglet-33Β 
posted an update about 1 month ago
view post
Post
3081
We are announcing the first generation of the Pebble model family!

These are the models we are releasing:

- Pebble 10M
- Pebble 25M
- Pebble 50M

Each model will use a Mamba-Transformer 3:1 hybrid architecture and will be pretrained on 25 billion tokens before IFT and SFT.

Depending on development time and resources, we may also release:

- Pebble 5M
- Pebble 75M
- Pebble 1M (possibly)

We hope you're excited and enjoy the models!

Follow for more:
@Hoglet-33
basically-ai
  • 7 replies
Β·
FlameF0XΒ 
posted an update 3 months ago
view post
Post
327
Hello, people of Hugging Face!

I recently released FlameF0X/TinyMoE-100m-2x8-retrained, a small Mixture of Experts language model trained on the Smollm-Corpus. Built on top of the Mixtral architecture, it’s fully compatible with πŸ€— Transformers right out of the box!

The model can produce somewhat coherent text on its own, and for some reason, it generates even more coherent responses when given a ChatLM template.

I’m excited to see what you all come up with, and feel free to fine-tune it if you’d like. In the meantime, I’ll be working on developing the chat-trained version.

Demo: FlameF0X/TinyMoE-Playground
Collection: https://huggingface.co/collections/FlameF0X/tinymoe