Full arm arrays are ready including a caption arm trained with danbooru, photography, and cc12m training.
These will be the primary test agency for the utilities and diffusion experimentation's introduction.
Full arm arrays are ready including a caption arm trained with danbooru, photography, and cc12m training.
These will be the primary test agency for the utilities and diffusion experimentation's introduction.
I've found peft style merging to be a continuity destroyer in many ways. I've been developing distillation methods for regularization techniques for this exact problem.
Standard PEFT LORA do not accurately account for the majority of LORA uses. Merge being a large problem of mine as I've made many constructs, and I can't simply create a LORA to merge to the next stage and decompose the differences later. The LORA and the actual model weights become interdependent, so when you remove the LORA space the trained space and behavior isn't available to actually attribute and extend. By consequence, the knowledge is often catastrophically forgotten within a few hundred steps.
I've made other systems but they can't be merged in correctly, so instead I've been refining forms of loss to allow multiple simultaneous loras to exist and to train alongside a model as a modular agency.
Direct merging is always going to create a continuity problem if you don't merge and test iteratively for data destruction. Essentially testing the output after each wave of integration, which is a clear overhead and time sink. It's worth it though, and will tell you which pieces of your loras you lost, and which pieces remain based on the exact expected behavior.
After that you can reinforce the good behavior and punish the bad behavior as per standard reinforcement. RNN could be employed to handle standard reinforcement as well.
Alright we snap the arms off and reconnect them at the end for post-training refinement. It's decided, the cost isn't turning a yield.
The kickover happens tonight at about 1 am, at the completion of stage 4. All four arms will be sidelined until the end of the trunk training completes.
It was a good experiment, that portion of the experiment ends now. We finish training the trunk without them, and then we build the collective of arms after.
Tests are showing the model is learning the same information as the arms, so they aren't cooperating as expected. I recall a multitude of experimental memory-based modules that would function more effectively for this exact system, so upcoming experiments for next week on the 3s variant will include those.
Instead of a collective they are forming an echo ensemble, which is the opposite of effective. Statistically the modules gain, while each subsequent stage introduces decay and destruction to the former. Given a few stages the model has already forgotten how to use the first arm.
Newly trained arms are done within 10 minutes rather than hindering the model training for days. This is a far faster method of experimentation. Alongside rapid updating the earlier arms happens as quickly as well, training new ones being considerably slower. The old arms are valuable utilities that cannot be disposed of.
Without the updated versions, the attached arms hinder the core model with the outdated arm information, becoming an active piece of information that cannot learn and adapt to upcoming information as effectively as required.
So they are to be temporarily removed, and those same arms retrained at the final stage, introducing new arms to be trained as well.
The arms themselves are important to a further experiment set, and I believe this result shows exactly what should always be expected when training a model's base along with the same information relayed into a divergent set of weights.
Those arms were meant to stay stubborn, keep the information learned during. Contradicting information causes catastrophic forgetting in some rows, complete forgetting in others depending on the severity. The continuity can't be easily measured, so the prudent course of action is to remove the variable from the experiment and continue.
For optimization, the differentiation to the information will be ignored if it's less optimal than the original, and the original is the optimal route. The more accurate is saved, and the trunk continues learning while the arms stay stubborn.
The optimal path will always be chosen unless the optimal path is not differentiated.
That's essentially the outcome with the arms, so we snap them off and continue the trunk to completion. With the finalized trunk we will have plenty of data to work with.
As of step 148,000~ the last arm linked trunk ends, and afterword the independent trunk continues, which I will begin experimenting on in different ways than currently experimented on.
The AMOE structure is about to get some experimental sidekicks.
The trunk should complete October 4th.
The multi-arm composites with the changes do not seem to have taken. The process likely needs to be halted and evaluated.
It's a bit too early to say, but it seems the loss function wasn't strong enough. The arms did not learn enough useful information.
As it stands, it seems the arms need a bit of a frozen kickstart to get going. Otherwise, the model never learns to utilize them for more effective information processing. The EASE of entry is harder for the arms to utilize, than simply defaulting to the trunk, so the branching system doesn't build the necessary directions immediately.
One of those, can't find the path because it's too complex of an entry sort of situations. I have a few ideas for how to guarantee the flood-gate entry, and as it stands there's new information for how these arms are to be trained as well.
I've run into this problem in the past. The model can't reach the point to recognize how much more effective the arm is at assisting the measure, or the arm itself is a hinderance to the process so it's simply omitted. It's the result of needing a process and a task from a model that isn't optimal, and the non-optimal route is simply being optimized out.
So, the model arms need more candy space, more attraction.
The arms aren't dead, they just need to get a little kickstart. The quieting algorithm isn't working effectively enough - which is a different problem.
The arm learning flood will happen with the right incentive.
The Beatrix V3 model is a bit past halfway done cooking give or take. Currently heading towards step 140,000.
If you load the model to experiment, ensure you have all the arms trained with the model active as well, otherwise the model's capacity will be hindered.