Quantization and the number of LLM parameters in practice – testing local models for writing prose (Q8 or 30B?) – part II

The second loop was no longer so much about “whether it can be done at all”, but rather about “how to arrange this so that the models have something to build from?”. After the first attempts, it became clear that simply putting the characters, the world, the conflict, and a few important motifs into the context was not enough. The model can generate an interesting passage, a well-sounding scene, or a convincing dialogue, but if it does not receive a stable foundation, it very quickly starts filling in the missing pieces in its own way.

That is why, in the second loop, the most important thing was discipline at the input stage. The information had to be organized, contradictions had to be eliminated, the relationships between events had to be clearly defined, and it was necessary to decide which elements of the world were truly binding. There was less faith in the idea that “the model will figure it out” and more work put into making sure it did not have to guess. In other words: instead of expecting the LLM to assemble loose notes into a coherent structure on its own, the material first had to be prepared in a way that made it possible to build that structure.

And although this is only the beginning of the article, I think it is already a good summary of working with LLMs today: the better the input, the better the output. It is not just about a nicer prompt, but about data quality, order in the assumptions, the absence of mutually contradictory information, and clear rules of operation. The model can help create, develop, and test ideas, but it will not replace order where there is chaos at the starting point.

Second creative loop – how to build a multi-query process?

An aspect that required particular attention was the model parameterization and quantization, and their selection at each stage of the loop. While theoretically, a model with more parameters but poorer quantization should be the better model:

I was tempted to use only models with the highest possible number of parameters at the expense of quantization (30B, or even 70B, which I tried in the first loop), but as I worked, I noticed that the best character depth and significantly more creative vocabulary were achieved with models with high quantization (Q6, and ideally Q8), even with fewer parameters. The conclusion is this: while a model may theoretically be “smarter,” it doesn’t necessarily mean it will be “better” for a given application, and sometimes better quantization will outweigh the higher number of parameters.

Additionally, I came across models fine-tuned for specific purposes—in this case, creative work. However, there are no rankings of models fine-tuned for writing and creativity, as these processes are very difficult and subjective to evaluate. However, after conducting approximately 50 tests, I identified my—obviously subjectively—best candidates. This allowed me to obtain a much better database. I also paid much more attention to the context size, as this is crucial for long texts.

These and other corrections combined resulted in a story of quite good quality, which I finally translated by hand – correcting the language, conversations, minor inconsistencies and giving the whole thing the so-called “final touch”.

Second loop – technical description of the test process

I knew that to achieve a better final result, I needed to give the LLM a better starting point. So, once I had an idea, I simply had to refine and clarify all the steps and elements, which, I hoped, would result in a better final result.

I also noticed longer and more precise model queries. Compared to previous commands, these were more extensive and more precise. However, I didn’t want to deliberately overcomplicate them so as not to limit the model’s creativity.

Moreover, for more creative tasks, I tested higher temperatures to obtain more creative answers, which overall had a positive impact on the final quality.

Temperatures range from 0.7 to 1.8 in some models. Different models react to temperature differently. Another interesting parameter was the “repeat penalty.” Increasing it generally reduced the number of repeats.

Second model list

In this selection, I was less concerned with “benchmark wisdom” and more with practical features: context, meaningful quantization, and whether the model had something conducive to writing. Therefore, instead of relying on models that performed best in logical tests, I used models that:

  • They had the biggest context
  • They had the most precise quantum possible, even at the expense of size
  • They were trained and specialized in writing (and therefore fine-tuned with datasets that improved these elements)
  • A potential additional advantage was the “abliteration” or “decensored” models – i.e. models with their security features removed, which was supposed to result in darker and more “graphic” descriptions, and I also thought that such models would cope better with “bad” characters
ModelContextSize and QuantumLicenseUse
Gemma-The-Writer-N-Restless-Quill-V28k10B Q8Apache-2.0Support basic
L3-DARKEST-PLANET8k16.5B Q8Apache-2.0Course events and prose
L3-DARKEST-PLANET-Seven-Rings-Of-DOOM8k16.5B Q8Apache-2.0Course events and prose
MN-DARKEST-UNIVERSE-29B131k29B Q6Apache-2.0The Course of Events and Prose
Rombos-Qwen2.5-Writer-32b-Q6_K32k32B Q6Apache-2.0Prose
Capybara Hermes32k7B Q6Apache-2.0Support basic

Necessary components

Inquiry for Capybara Hermes and Gemma The Writer:

Assume the role of creative writer helper. Your task is to help writer to create a coherent and interesting world. Propose what elements should be defined before starting writing a story and what those should contain (for example: background history of world where action takes place).

Because the Gemma and Capybara models are more creative, the answers were written in flowery language rather than informative, but that didn’t hinder me in any way. Despite this, I didn’t receive any suggestions for additional elements, so I continued working with the same list, which I didn’t mind, as I think this base should be sufficient. However, I did receive some interesting suggestions for describing the world, which I decided to incorporate at a later stage. The compiled list looked like this:

  1. The subject of the work – a description of the general setting of the world
  2. Description the world presented
    1. Government and politics
    2. Economic base (technology and resources)
  3. Safety data sheets characters
  4. Literary style and recipient
Gemma The Writer’s model is more creative than the models on the first list.
Capybara Hermes also gave pretty good answers

Corrections to the main thread

So far, I have used the following description for AI queries:

In small medieval town there is a plague of mysterious robberies and thefts. Local sheriff attempts to use “Justice System” – mysterious object analyzing diagrams – to detect thefts. He is assisted by his sidekick, who does not really trust the Justice System.

From the perspective of the previous answers and the result of the first loop, such an important element used in the context of LLMs should have a more precise description almost everywhere and already include an outline of the plot, so that all other generated elements would take it into account. Since I already had an outline of the plot, I could expand on this description significantly.

In small medieval town Kopperwah there is an ongoing plague of robberies and thefts. Local Sheriff Jameson attempts to use “Justice System” – a software run on his sidekick AI robot named Gideon – which analyzes schemas to detect thefts. Because for two years there is no progress, king sent his agent – Erica – to help with this case. At the same time villagers are trying to use Heartstone for the same purpose, but without success so far.

This main plot definition clearly defined what I wanted to achieve and improved the results of subsequent queries. For subsequent queries, I used this new topic description wherever helpful. This meant that, regardless of which element of the story I was building, the single key truth of the underlying plot description was almost always retained.

Description of the world presented

I wanted the description of the world to be more precise and somewhat broader – in other words, I wanted the world to be not only “pretty” but also functional: politics, economics, and technology would begin to explain why something bad might happen at all and why it matters. I chose the Qwen Writer model for this because, in testing, it demonstrated a good balance between creativity and instructional execution. He was tasked with creating a much more precise description of the world, which he accomplished admirably, producing several paragraphs. I also wanted the model to be as creative as possible, so I used a temperature of 0.9 and set the maximum number of response tokens to 6096.

Settings for Rombos Qwen Writer – what’s great about GPT4All for beginners is that the settings are described, making it easier to experiment with them

In response, I received a decent description of the world, which I then manually edited for use in future queries. It’s worth noting that I completely removed the characters from this list, as they will be defined in detail in their data sheets anyway. The final result looks like this:

1. History: Kopperwah was founded 80 years ago by settlers seeking refuge from the chaos of war-torn kingdoms to the east and north. The town prospered due to its strategic location on neutral area along trade routes connecting these rival realms. Known for its good grade copper and trustworthy merchants. In recent years an increase in crime has plagued Kopperwah , threatening both safety and prosperity.

2. Major events: Two significant occurrences have shaped modern-day Kopperwah :

– Finding copper and building mines: 75 years ago

– Grant town status from king Aldric III and establishing peaceful trade relations with neighboring towns: 60 years ago

– Great Plague which decimated the population and greatly reduced trade: 10 years ago

– Heartstone discovery in copper mines: 2 years ago

3. Geopolitical landscape: Kopperwah is an autonomous town within a larger kingdom ruled by King Aldric IV. technical While under royal jurisdiction, it enjoys considerable self-governance due to its economic importance as a trading hub. Tensions occasionally arise between the king’s agents and local authorities over matters of law enforcement and taxes. Aldric III and Aldric IV being intelligent kings learned and earned a lot by observing fighting kingdoms in the north and east.

4. Cities and settlements: Nearby towns include:

– Rolleypick (to the north, behind the border): A community of farmers and fishermen.

– Willowbrook (to the east, behind the border): Home to skilled blacksmiths and armor makers.

– Ravenhold (to the south, within kingdom borders): Renowned for its herbalists and healers.

5. Geography and climate: Kopperwah is nestled in a valley surrounded by rolling hills, with mild temperatures year-round due to its proximity to both mountains and forests. The town sits along the banks of the Copperstream River, which serves as an important trade route connecting it with Rolleypick and Willowbrook.

6. Economy: Kopperwah’s economy is based on copper mining and trade with adjacent towns. Kopperwah earns a lot of money as a neutral trade hub between fighting kingdoms of north and east, facilitated through its central location between rival kingdoms. Local artisans produce high-quality goods such as textiles, pottery, and copper art that are traded extensively throughout the region. The town also benefits from being a waypoint for caravans transporting valuable resources like spices, silks, and precious metals.

7. Cultural context: Kopperwah is known for its vibrant festivals celebrating seasonal changes and harvests. Adjacent towns of Relleypick and Willowbrook often ignore war and arrive to Kopperwah to rest from war and have some fun. Residents place great value on community cooperation and mutual support during times of hardship. Despite the recent surge in crime, most villagers maintain a sense of optimism about overcoming challenges together.

8. Technological infrastructure: wartime technology is heavily based on copper – which makes Kopperwah strategically a very important town. But since the town is outnumbered by plague it became easier to target by thieves. In response to rising criminal activity, Sheriff Jameson has implemented innovative technological solutions like the Justice System – exotic software used by Gideon (robot and his companion) designed to analyze patterns and predict potential thefts or robberies before they occur. These advancements represent cutting-edge developments in law enforcement within the kingdom.

9. Unique features:

– Heartstone : An ancient artifact said to have been crafted by powerful mages long ago, imbued with the ability to reveal truth and falsehood based on hearing the person speak. Only responds to Gynerva.

This world description, as well as the main thread, was used in subsequent queries as part of the “system prompt”—the model’s underlying instructions. This might seem like a lot for an instruction, but it’s not. It’s a quantity that even locally run models can handle without any problems.

Character Data Sheets

At this point, I began to sense that a single paragraph per character would no longer suffice, and the potential conflicts of interest between the characters became increasingly apparent. To avoid hallucinations, I decided to expand characterizations with elements such as physical descriptions, personality traits, skills, motivations, short backstories, connections to other characters, and hidden motivations. While not all of these elements would necessarily be directly used in the story (just as not all elements of the setting), they could indirectly influence the characters’ actions and therefore needed to be included in context.

I used the Qwen Writer model again because it did a good job of describing the world. For each character, I ran a separate query using the following formula:

Assume the role of creative writer helper, tasked with describing character for story. You should be very creative and descriptive. Character should be compatible with World Description and fit into this world. It’s characteristics should be clearly outlined and affect their actions within the described World.

Character Card should have the following :

1. The name of the character.

2. Physical Description (weight, height)

3. Personality Traits (strengths, weaknesses)

4. Motivations: What drives the character? Their goals, desires, fears, and values.

5. Backstory: A brief summary of the character’s past experiences, including childhood, significant events or relationships, traumatic experiences (if applicable)

6. Natural abilities or skills they possess, formal education or training received

7.Key connections to other characters, including family members, friends and allies, enemies or rivals

8. Goals & Conflicts: What the character wants to achieve (short-term and long-term) and what obstacles they face in pursuing those goals.

9. Unique mannerisms, speech patterns, or behaviors that make them stand out

10. Motivational Quotes: Inspirational quotes or phrases that drive the character’s actions or decisions.

Describe character:

[Description of the characters from the first loop]

Due to the increased creativity of the models, some inconsistencies started to appear, so I had to correct these cards manually (mixed place of birth, description of Gideon as a human, or information that I didn’t plan to use and would only make further work harder for the LLMs).

The result of this work was lengthy character sheets—almost 19,000 characters in total—and that was without any description of the world. At this point, I suspected the story would be very dense, so as to fit in as many facts as possible, both about the world and the characters.

Literary style and audience

I added a bit of flair here because I figured if I was going to test uncensored models, at least there should be a real difference in the tone and energy of the scenes. I wanted the uncensored models to be able to introduce a bit of brutality and vulgar language, so I adjusted the prompt to be more “naughty.” As a result , was created following instructions :

Story should be written for well-red adults that like science-fiction and fantasy, with occasional usage of archaic languages. You can use: graphic descriptions including blood, a bit gore, violence, strong language and strong interactions between characters.

Such a description may seem too strong, but it must be remembered that models tend to average – and in this case “shallow” – instructions, so the result was not drastic at all, but rather at the “teenager” level.

The course of events

This part of the work was quite “technical,” because a coherent plan of events is something between logic and drama for the models—and this is where their confidence ends and improvisation begins. RAG still struggled because it lost too much detail. I’m not sure why—perhaps it’s due to the very nature of RAG, perhaps a different engine or different prompts were needed. Overall, to achieve the greatest possible consistency, I decided to contextualize everything. I tested several configurations, but the standard approach was to combine all the above elements. So the example prompt looked like yes :

You are a helper of a professional writer. You are given in attachments elements of the world: topic description, world description, character cards and writing style. Your task is to generate a list of events for a short story.

Story should be written for well-red adults that like science-fiction and fantasy, with occasional usage of archaic languages. You can use: graphic descriptions including blood, a bit gore, violence, strong language and strong interactions between characters.

In small medieval town Kopperwah there is an ongoing plague of robberies and thefts. Local Sheriff Jameson attempts to use “Justice System” – a software run on his sidekick AI robot named Gideon – which analyzes schemas to detect thefts. Because for two years there is no progress, king sent his agent – Erica – to help with this case. At the same time villagers are trying to use Heartstone for the same purpose, but without success so far.

And below that there was a description of the world presented and all the character cards.

I tested the prompt on all the models from the “second model list,” and the result was that I wasn’t satisfied. The response from Gemma Writer was utter gibberish, while the other responses were unsatisfactory—the timelines weren’t detailed enough, were inconsistent, and there wasn’t enough action between the characters. They even suggested that the sheriff was incompetent, failing to employ even basic investigative techniques (like “asking the townspeople for help in gathering information”) for two years. On the other hand, there were some good ideas (like Sigils) that I gladly implemented.

Darkest Planet – a pretty cool description of the events
Seven Rings – also using context instead of RAG

Ultimately, none of the answers were fully applicable, so I used them all as a basis for creating my own timeline. I was still about 15% short and couldn’t come up with anything (at about 70% of the story) because the events suddenly seemed “too dense.” The solution was to ask Gemma Writer to fill in the missing pieces. As a result, I had a pretty good timeline, ready for my final creative process.

Finishing the story

By this stage, I felt like “it was roughly working,” but I still needed to structure the process so the models wouldn’t diverge in style, facts, and tone from one day to the next. Why “one day to the next”? Because instead of having the model write everything at once, I’d broken the timeline into days, which created a natural, consistent rhythm to the novel and helped the models avoid getting lost.

From there, things were generally simple. This time, I provided the entire event schedule and world description for the LLMs, and the LLM was tasked with creating ONLY one point (day) from the event schedule . looked as follows :

You are a helper of a professional writer. You are given in attachment world description.

Story should be written for well-red adults that like science-fiction and fantasy, with occasional usage of archaic languages. You can use: graphic descriptions including blood, a bit gore, violence, strong language and strong interactions between characters.

Main topic is: In small medieval town Kopperwah there is an ongoing plague of robberies and thefts. Local Sheriff Jameson attempts to use “Justice System” – a software run on his sidekick AI robot named Gideon – which analyzes schemas to detect thefts. Because for two years there is no progress, king Aldric IV sent his agent – Erica – to help with this case. At the same time Gynerva and Eleanor are trying to use Heartstone for the same purpose, but without success so far.

Your task is to generate next section based on the plan.

Events plan:

[plan of events pasted here]

Generate story for day number [X].

Another experiment was to alternate between several LLMs that I found worked best—Darkest Planet, Qwen Writer, and Darkest Universe—to add variety and dynamics to the writing style.

One chat for each “day” of the event description.

At the very end, of course, I simply combined all the days into one string and translated the result manually from English to Polish, making only very minor corrections.

I’m aware of the many literary shortcomings of this result, but one of my goals was to interfere as little as possible with the final result. I wanted the AI to be the main author of the story. Therefore, the final version is clunky and has flaws that could certainly be corrected by hand or with larger cloud models – but that would be inconsistent with the original intent, so I didn’t.

What can writing a story teach an engineer?

I am leaving this fragment as honestly as possible, because it is this “cold” way that best shows what one can take away from this experiment.

I’m happy with the story itself, but it seems too intense, and at the same time, a lot of material wasn’t used. I feel like I could make a small book out of it, but the sheer scale of the local contexts is a significant limitation.

Overall, you can learn a lot by repeating such an experiment. For example, you can:

  • learn a lot about working with LLMs
  • get to know some of their “natural” limitations
  • learn to work with context and the effects of transcending it
  • adjust “creative” parameters
  • get to know a lot of interesting models
  • better understand the differences between quantization and the number of parameters
  • improve your prompt engineering skills

And, as a bonus, you can generate a decent story. Therefore, I would recommend this – seemingly non-technical – exercise even to those who work with highly technical models on a daily basis.

Of the conclusions that can be drawn from such an exercise, I think the most crucial will be these:

  • Sometimes, better quantization is worth more than the number of parameters. In the case of prose, it allowed for a richer vocabulary and deeper relationships between characters.
  • Working with RAGs may not be as straightforward as intended, and it’s easy to miss important pieces. While current RAG systems are far superior, the model must be able to communicate effectively with such data.
  • Context and messaging are crucial when it comes to “conversations” with models. They should always be verified, especially for specific queries.
  • A model trained specifically for a given task will perform best. Specialized models (for writing, coding, or research) will perform better than general models in a given domain. If a ready-made model isn’t available, sometimes you need to train it yourself.

An additional conclusion is that LLMs can be a great aid in creating stories—for example, when you’re stuck or don’t really know where to begin. However, achieving character depth, interesting relationships between characters, and a solid plot twist is difficult. Furthermore, consistency is key. It’s also possible to “talk” to characters by formatting their responses appropriately, but you often need to know in advance what you want to achieve.

If you’d like Sailing Byte to help you select a model, train it, and prompt-engineer it to achieve the best results for your goals , please contact us using the form below. We can create a SaaS solution with a dedicated AI model that’s optimal for your application.

The key is that the experience gained from such an exercise can be transferred to other fields, and as a result, you can simply become a better “prompt engineer.” So, it’s worth trying such an exercise, locally and for free. And while you shouldn’t expect miracles at this stage of AI development, in the end, it’s also just good fun with models.

Author

Łukasz Pawłowski

CEO of Sailing Byte

Sailing Byte CEO and former PHP developer. Founder of a software house specializing in a partnership-driven approach, with expertise in Laravel, React.js, and Flutter. My objective is to deliver scalable SaaS solutions through Agile methodologies—offering clients a blend of experience, knowledge, and the right set of collaborative tools. To achieve this, I am committed to sharing my expertise on this blog with clients and readers across Europe, the UK, and the USA, empowering their businesses to flourish.