Full arm arrays are ready including a caption arm trained with danbooru, photography, and cc12m training.
These will be the primary test agency for the utilities and diffusion experimentation's introduction.
Full arm arrays are ready including a caption arm trained with danbooru, photography, and cc12m training.
These will be the primary test agency for the utilities and diffusion experimentation's introduction.
I've found peft style merging to be a continuity destroyer in many ways. I've been developing distillation methods for regularization techniques for this exact problem.
Standard PEFT LORA do not accurately account for the majority of LORA uses. Merge being a large problem of mine as I've made many constructs, and I can't simply create a LORA to merge to the next stage and decompose the differences later. The LORA and the actual model weights become interdependent, so when you remove the LORA space the trained space and behavior isn't available to actually attribute and extend. By consequence, the knowledge is often catastrophically forgotten within a few hundred steps.
I've made other systems but they can't be merged in correctly, so instead I've been refining forms of loss to allow multiple simultaneous loras to exist and to train alongside a model as a modular agency.
Direct merging is always going to create a continuity problem if you don't merge and test iteratively for data destruction. Essentially testing the output after each wave of integration, which is a clear overhead and time sink. It's worth it though, and will tell you which pieces of your loras you lost, and which pieces remain based on the exact expected behavior.
After that you can reinforce the good behavior and punish the bad behavior as per standard reinforcement. RNN could be employed to handle standard reinforcement as well.
Alright we snap the arms off and reconnect them at the end for post-training refinement. It's decided, the cost isn't turning a yield.
The kickover happens tonight at about 1 am, at the completion of stage 4. All four arms will be sidelined until the end of the trunk training completes.
It was a good experiment, that portion of the experiment ends now. We finish training the trunk without them, and then we build the collective of arms after.
Tests are showing the model is learning the same information as the arms, so they aren't cooperating as expected. I recall a multitude of experimental memory-based modules that would function more effectively for this exact system, so upcoming experiments for next week on the 3s variant will include those.
Instead of a collective they are forming an echo ensemble, which is the opposite of effective. Statistically the modules gain, while each subsequent stage introduces decay and destruction to the former. Given a few stages the model has already forgotten how to use the first arm.
Newly trained arms are done within 10 minutes rather than hindering the model training for days. This is a far faster method of experimentation. Alongside rapid updating the earlier arms happens as quickly as well, training new ones being considerably slower. The old arms are valuable utilities that cannot be disposed of.
Without the updated versions, the attached arms hinder the core model with the outdated arm information, becoming an active piece of information that cannot learn and adapt to upcoming information as effectively as required.
So they are to be temporarily removed, and those same arms retrained at the final stage, introducing new arms to be trained as well.
The arms themselves are important to a further experiment set, and I believe this result shows exactly what should always be expected when training a model's base along with the same information relayed into a divergent set of weights.
Those arms were meant to stay stubborn, keep the information learned during. Contradicting information causes catastrophic forgetting in some rows, complete forgetting in others depending on the severity. The continuity can't be easily measured, so the prudent course of action is to remove the variable from the experiment and continue.
For optimization, the differentiation to the information will be ignored if it's less optimal than the original, and the original is the optimal route. The more accurate is saved, and the trunk continues learning while the arms stay stubborn.
The optimal path will always be chosen unless the optimal path is not differentiated.
That's essentially the outcome with the arms, so we snap them off and continue the trunk to completion. With the finalized trunk we will have plenty of data to work with.
As of step 148,000~ the last arm linked trunk ends, and afterword the independent trunk continues, which I will begin experimenting on in different ways than currently experimented on.
The AMOE structure is about to get some experimental sidekicks.
The trunk should complete October 4th.
The multi-arm composites with the changes do not seem to have taken. The process likely needs to be halted and evaluated.
It's a bit too early to say, but it seems the loss function wasn't strong enough. The arms did not learn enough useful information.
As it stands, it seems the arms need a bit of a frozen kickstart to get going. Otherwise, the model never learns to utilize them for more effective information processing. The EASE of entry is harder for the arms to utilize, than simply defaulting to the trunk, so the branching system doesn't build the necessary directions immediately.
One of those, can't find the path because it's too complex of an entry sort of situations. I have a few ideas for how to guarantee the flood-gate entry, and as it stands there's new information for how these arms are to be trained as well.
I've run into this problem in the past. The model can't reach the point to recognize how much more effective the arm is at assisting the measure, or the arm itself is a hinderance to the process so it's simply omitted. It's the result of needing a process and a task from a model that isn't optimal, and the non-optimal route is simply being optimized out.
So, the model arms need more candy space, more attraction.
The arms aren't dead, they just need to get a little kickstart. The quieting algorithm isn't working effectively enough - which is a different problem.
The arm learning flood will happen with the right incentive.
The Beatrix V3 model is a bit past halfway done cooking give or take. Currently heading towards step 140,000.
If you load the model to experiment, ensure you have all the arms trained with the model active as well, otherwise the model's capacity will be hindered.
The deeper variant is showing quite weak recall in comparison to the v1, 2.5s, and 2s softmax before softmax collapse.
There's a definite problem here that needs to be addressed.
Deep byte recall wasn't exactly filling the 4096 space for the 2s, while the 2048 space was predominantly filled by v1 give or take. The models never quite coalesced, which v3 was expected to fill the space for.
Instead of filling the space as expected, the model is showing weaker code recall later in the chain. I'm not going to pull the plug yet, but the results this far into training should be much cleaner than they currently are.
The 3s model has been showing improvement, so I'm going to be cautiously optimistic. I'm thinking this model will need substantially more training than the softmax version, which means something needs to change in the attention mechanism to converge more quickly.
The 60b tokens very well might not be enough, which isn't a good sign. This larger investment is definitely going to need to be cut down to size if I want reasonable prototypes.
The 2s variant showed improvement throughout, and the end result of the 2s variant wasn't as strong as the v1 mixed but it was quite strong. The 2s full softmax showed very strong curves before the collapse, which signals to me we need a late-stage mechanism control for softmax, rather than a full replacement for softmax. The mixed attention are likely the best course of action. Softmax + splat, potentially a new mechanism to replace splat to ensure the delta isn't wildly misappropriated or drifts so far that the model cannot recall what was just consumed.
Verdict: Run will complete between October 14th and October 17th. A bit pushed out, but considerably less time than the alternative.
I've settled on a compromise. It will yield a substantially weaker abstentation policy, 1/4th of the cells. Without the necessary hardware involved I must weaken the model, but the results themselves will not be fruitless.
After the run is complete, I'll run multiple tests to determine the BEST possible mix for version 4. The BEST possible compromise. The STRONGEST possible yield that builds the most effective information, without destroying the necessary yield and information.
As it stands, the model will yield. The results will be potent. With any luck, they will be more than potent enough for distillation, diffusion, classification, and more experiments.
With that compromise, I'll be attaching additional curve watchers for the risk. It's a minimal risk, but it's high enough to put some gauges on the model. The cost for this model isn't cheap.
Seems we've run into engineering problems. The time degraded from 4 seconds per step to 9 seconds per step, compounding increase per independent arm attached.
I need to halt and run analysis. As it stands, we'll run until the end of Friday, then the engineering goes up on blocks.
I'm going to squeeze every bit of every byte from this engineering. Formulas up for scrutiny, system optimizations up for scrutiny, kernel optimizations, libraries, everything. Running until October 21st isn't reasonable. The models must be prepared correctly and reasonably to the mathematics, the results must be calculated correctly and the model trained to parity without faulty modifications during training that can impact the results.
Here's an interactive viewer for the internals of Mini-Beatrix-2.5s
I'll enhance it for 3 when it's ready.
As a comparison to global erank, we're looking at a structure of 400+ for around half of Beatrix V3 so far, so roughly 16+ blocks of erank >400, substantially stronger than the original two models for geometric attribution. The final block has a collapsing problem currently, but I believe others have the answer with autoregression models through a finalized projection smoothing layer concept. I haven't employed it yet though.
The fractal instability hits pretty early. You need rounding structures early otherwise the gradients explode at one point or another. The predominant problem was loss explosions. It happened because of ill-formed eigens in the intentional step structure I was experimenting with. 5 step cantor essentially ensured the cantor fractals deviate to a certain degree, and depth itself was meant to raise the steps of fractals to new states and interpolate the fractals.
If you use any of this, make sure you either pass it into AI for optimization - as it's likely terrible due to being my earlier models (I came from game development, optimization is very different). AI will be able to improve the speed and accuracy of the formulas.
https://github.com/AbstractEyes/geofractal/blob/main/src/geofractal/model/core/vit_beatrix.py
https://github.com/AbstractEyes/geofractal/blob/main/src/geofractal/model/positional/cantor.py
https://github.com/AbstractEyes/geofractal/blob/main/src/geofractal/model/core/geo_fractal_david.py
One of the problems was similarity. Almost everything was self similar, which in theory should have helped differentiate. However in practice, the structure found it's own similarity attractor basins that cause cascade corruption down the chain. The only solidity was to introduce eigen comparators through decomposition learning, which is a little different than autoregression. With this, the decomposition required more accuracy otherwise the system would always default to 1 of the first 3 steps - resulting in rigid or slightly less rigid articulations.
I measured fp64 being required for stable 4 step, and fp roughly 92 to be in a safe zone for stable Mandels at step 5. Julia requires something substantially larger than mandels. Fp64 is ENOUGH for rotary offset in standard positional systems, however fp128 is required for something akin to cantor fractal positional systems of differentiation.
It happens due to the eigenvalues themselves often malforming, and the subsystem silently rounds them. Using FULL SVD is a compositional fix for comparison, with that introduces a huge overhead as well.
Fractals themselves turned out to be more compositionally useful, not as additive elements, but as miniature rounding structures. The splat there was built under the concept of eigen substitution, meant to composite a series of tiny opinions from tons of subsystem residuals together into a composite "blackboard", forming a more robust and structural aligned INK BLOT splat, similar conceptually to viewing a random inkblot. This eventually composites into a utility of structural awareness, and it really doesn't take very long.
Essentially, that structure is geometric in nature, but it's not using Eigenvalues directly. It CAN use them, it should be capable of using any structural bounds with attributable contributions.
Splat functions viably at bf16, is a bit slower than MHA, but houses geometry more cleanly than MHA (sometimes by a huge margin) when trained with MUON instead of adam, adamw, or another multitude of optimizers I ran. I have attempted custom optimizers to encourage this behavior further, but the results showed MUON is just better.
Give it a shot in something simple, it'll train fast enough.
Pretty much anything in here is useful.
https://huggingface.co/collections/AbstractPhil/geolip-research-concepts
Eigens and causal chains have correlations but not causation without additional contributions to the assessments, the SVAE shows this to be a guarantee in many shapes, and in many others impossible.
The accuracy between the two requires a smoothing system, alpha differentiation through projected MHA-esque alpha attention to patchworks in order to fill the gaps. They don't directly line up quickly though, it looks more soupy when it's done.
They coalesce, but the extractions aren't consistent enough to directly use without a series of wrappers and structural alignment systems. Cantor Aleph and Omegas are essentially this structural system, but they are unstable. Cantor fractals remain unstable until around fp128 for Mandelbrot without redefining the underlying methods the mathematics linalg system uses. I did some headway on this, but I ran into a glacier that I would have needed to sink months into to make headway so I built a system to replace the slower linalg systems and the system lost much of it's cantor fractal capacity in favor of reproducibility and consistency.
The prototype forged from a 52,000 battery sweeps to find the most consistent recon convergence over time, heavily scrutinized and analyzed for over a month to build into something useful.
https://huggingface.co/AbstractPhil/geolip-SVAE
The current best case of the eigens research conclusions. Everything SVAE built to the attention prototype, everything constellation built to the processing and lookups for the model, everything distillation built to the banks for capturing and yielding, everything structural built from the knowledge and wisdom of other researchers. Well, not everything structural - it needed a lot of geometric formation and structural cohesion to make it work.
https://huggingface.co/AbstractPhil/aleph-splat-0/blob/main/splat_attention.py
Attempts to speed the SVD up were somewhat fruitful, somewhat not. They are good for inference, but I never programmed the gradient backprops for it.
https://huggingface.co/AbstractPhil/svd-triton
The mobius lens being a faster form wasn't strong enough as an activation system. I needed an architecture around it, not just an activation.
https://huggingface.co/AbstractPhil/mobiusnet-distillations/blob/main/make_chart_1.py